Your site sends the text before publishing it and acts on the answer: publish it, hold it for a person, or refuse it. With ToxicFilter that is one POST to /v1/text with the content. You can add the languages you expect, the surface the text will appear on (a comment, a listing, a profile), the policy to apply and your own reference for it. The answer carries an id that names that one decision, for the review queue, a report of a false positive or a question to support.
Fifteen categories, not one toxicity score
Spam, toxicity, harassment, hate, sexual content, violence, self-harm, the safety of minors, scams, personal data, disguised text, gibberish, off-topic content, prompt injection and your own word lists. Each one has a review line and a block line. Some never block by default: toxicity only holds, because a swear word is not abuse, and self-harm only holds, because the person hurt by deleting that post is the one who wrote it. Subjects such as gambling or crypto are measured on a separate axis and act on nothing until a rule says so.
Disguised insults and leetspeak
A filter that matches the raw text only catches people who were not trying. Here nothing is matched raw: the text is folded first, so look-alike letters, digits used as letters, invisible characters, censor asterisks, accents and stretched letters all come back to the word they stand for. How much a word had to be unmasked is measured on its own, under evasion. Allowlisted words are blanked out before searching, so Scunthorpe, assess and cockpit are never flagged.
Fast on the free path, the model only when it adds something
Nearly all comments are somebody writing an ordinary sentence, and paying a model to read each one is how moderation ends up costing more than the site it protects. The free checks always run, and the model only reads when they left the question open. A check costs one credit; when the model reads it, the tokens it used are added, about eight credits in all for a comment. If the model provider is down, the answer says degraded instead of passing the comment off as read.
Rules per account or per call
Start with the shipped lines, or one of the policy templates for a community, a marketplace or a site for children. A stored policy is versioned, and a change can run in shadow beside the one in force, so you see what it would have done to your own traffic before it decides anything. When you would rather not configure anything, send rules in the call: "block sexual content from 0.7" is one request, and categories you do not mention decide nothing.
The word lists and phrase patterns cover English, Spanish, Portuguese, French, Italian, German, Catalan and Dutch. Tell the API which languages you expect, and a text in another one, or full of another alphabet, is reported as such. The model reads far more languages, and effort: high brings it in for a text outside the eight.