Call the API before a message, a comment or a post is published, and act on one of three answers: publish it, hold it for a person, or refuse it. Most of what pupils write is a question about homework or a reply to a classmate, and the free checks settle that in about a millisecond. The model only reads what they leave open. What lands in review waits in a queue, in your dashboard or through the API, and a signed webhook tells your platform when a teacher or a moderator decides.
The children template moves the lines a general community uses. It refuses harassment, hate, threats and sexual content sooner, holds profanity for a look without ever refusing a message for that alone, holds personal data sooner, and puts posts about self-harm in front of a person at a far lower score. Every number is yours to change, and any change can run first as a second policy beside the one in force, so you see what it would have done to your own traffic before it acts.
Bullying between pupils
An insult aimed at the reader is scored under harassment; the same word about a test or a game is not. Insults spelled to get past a word list (l0ser, a letter from another alphabet, a space in the middle) are folded back into plain letters before anything is matched. When several pupils turn on one in the same thread, send it to /v1/conversation: the number of different people being hostile becomes part of the verdict.
Protecting children in chats and messages
An approach to a child is not one message, so ToxicFilter only looks for it across a conversation, and it follows strict rules. A stated age is never a finding. Only what other participants say about their age counts, never the child's own words. The reasons describe shapes, such as asking for secrecy or a move to another app, and never what anybody is. A hit beside a stated child is refused and is a person's job, immediately. The instant checks read these shapes in eight languages, and with effort: high the model reads the whole conversation for what the wording leaves to context.
Personal data a child shares: mask it
A pupil who writes their phone number to arrange a group project has not done anything wrong. With redact, the answer carries the same text with the phone, the email or the card masked, and the rest can be published. Your school's name, or anything else specific to your platform, goes in a rule of your own words, and from then on it is found like any other.
Self-harm: hold, never delete
A pupil who says they want to die may be asking for help, and removing the post removes them. self_harm holds for review and never blocks unless you decide otherwise, the answer says why in words, and the moderation.review webhook can alert whoever on your side can respond. What pupils write is not stored: a verdict keeps a fingerprint of the content, never the content itself.