All guides
    Safety
    6 min read

    When something gets through the filter

    Every filter has a miss rate above zero. What separates schools is not whether something gets through — it is what happens in the following hour.

    Any vendor implying their filter is perfect is either misinformed or hoping you are. Content moderation involves judgement calls at a boundary, made at scale, by systems that are right most of the time. Most of the time is not all of the time.

    So the responsible thing is to decide in advance what happens on the day something gets through, when the people involved are upset and the temptation to react quickly is strongest.

    In the first hour

    1. Attend to the student first. Everything else can wait ten minutes; a child who has seen something upsetting cannot.
    2. Preserve the evidence before anything changes it. Screenshot what was seen, and note the account, the approximate time and the tool used. Do not delete the conversation — it is the only record of what actually happened.
    3. Do not interrogate. What was asked matters, but a child who feels accused stops being accurate, and you need accuracy more than you need a confession.
    4. Tell one designated person. Whoever holds safeguarding should know before the story spreads through the staff room.

    Then work out which kind of failure it was

    This determines everything afterwards, and the three cases look identical at first.

    A judgement call at the boundary

    The system evaluated it and allowed it, and reasonable people might disagree. This is the most common case and the least alarming. The remedy is to tighten the specific rule.

    A gap in coverage

    The rule existed but did not apply to this tool — most often, a topic blocked in chat that was not blocked in image generation. This is a product defect and worth raising sharply with the vendor, because it means the school's policy was never actually in force where it mattered.

    The check did not run

    The safety system was unavailable and the request went through unchecked. If a product can do this, it will do it again, and no amount of rule-tightening helps. Ask the vendor directly whether their system fails open, and treat a vague answer as a yes.

    Telling parents

    Tell them, and tell them before they hear it elsewhere. A school that reports its own incident is trusted; a school that is discovered not to have is not, and the second outcome is far more expensive than the first.

    Say what happened, what the child saw, what you have changed, and what you are asking of them. Resist the urge to minimise — parents are far more forgiving of an imperfect filter honestly reported than of a confident reassurance that later turns out to have been wrong.

    What to change afterwards

    Add the specific topic to your blocked list, in the words that actually came up rather than a euphemism. Check the same request in every other tool, not only the one where it happened. Ask the vendor what changed on their side, and whether other schools saw it.

    Then leave the rest alone. The instinct after an incident is to restrict everything, and a filter tightened into uselessness gets worked around, which leaves you worse off than before.

    The honest expectation to set

    When a school adopts a filtered AI tool, say to staff and parents at the outset that the filter reduces exposure and does not eliminate it, and that there is a plan for the exception. A school that has said this in advance is in a completely different position on the day than one that promised the problem could not occur.

    Ask us what Navōn does when the safety check itself is unavailable.

    Common questions

    What should a school do if inappropriate content reaches a student through an AI tool?

    Attend to the student first, preserve the evidence before it is deleted, avoid interrogating, and inform the designated safeguarding person. Then determine whether it was a boundary judgement, a gap in tool coverage, or the safety check failing to run.

    Can any AI content filter be 100% effective?

    No. Moderation involves judgement at a boundary at scale, so the miss rate is above zero. A vendor claiming perfection is not describing a real system.

    Should we tell parents when something gets through?

    Yes, and before they hear it elsewhere. Parents forgive an imperfect filter honestly reported far more readily than a reassurance that later proves wrong.

    Published by Navōn. How these guides are written and checked.

    Read next