Why one toxicity score is not enough
A gaming forum and a children's platform should not refuse the same sentences. A single number with a single cut-off bakes one site's tolerance into every other site. ToxicFilter scores fifteen categories separately, plus subjects (gambling, crypto, counterfeits) and lead types on their own axes, and each one acts only at the line a policy gives it. The two examples above that change decision are the same text scored the same way: only the lines moved.
How a moderation policy works
A policy stores what you changed and nothing more. The categories you never touched keep following the shipped numbers, so when those improve, the change reaches you too. Switching a category off is a line above 1: it keeps scoring and keeps appearing in the answer, it never acts. Your word lists sit beside the lines: words to refuse, words to look at, and words never to flag, which are blanked out of the text before our own list searches it. A description of what your business does helps the model judge what is off topic for you. Lines can also differ by surface, so a profile and a comment can be held to different numbers under one policy.
Templates per kind of business
Creating a policy starts with what it is for: a community or game, a marketplace, a contact form, a dating app, classifieds, a job board, reviews, an AI product or a platform for children. The template's lines are copied in at that moment and are ordinary rules from then on. Nothing is looked up again later, so editing your policy never surprises anybody else's, and a change to the template does not rewrite yours.
Shadow mode: test a policy on real traffic
The question before every threshold change is what it would have done to last week. Set a second policy as the shadow of the live one. Every call is judged under both, the live one decides alone, and the record keeps both decisions, so the dashboard can show the disagreement over a period of your own traffic. The model's findings are reused as they are and only the free checks run again, so a trial does not double the bill. Images are not re-run, and nothing else changes: same answer, same billing, no extra webhook.
Versions, records and rules in the call
Saving a policy makes a new version. The slug and version come back in every answer and are copied into every record, and the version is part of the cache key, so a verdict reached under the old rules is never served again. A name that does not exist is a 422, never a silent fallback. When you would rather not configure anything first, rules in the request carries the lines for that call alone; it is recorded as inline, which is why a decision that has to be defended later belongs in a named policy.
Reputation, kept on a short leash
A policy can let a person's own record in the same project move the lines a little: at most 0.10, only after 20 verdicts, never across accounts, always reported with the adjustment. It never switches a line on or off, and a good record never loosens minor safety, self-harm, violence or hate, because earning trust first is exactly the pattern those categories exist to see through.