How to moderate comments and posts in real time
Call the API before a post, a comment or a message is published, and act on one of three answers: publish it, hold it for a person, or refuse it. Most of a community's traffic is someone saying something ordinary to someone else, and the free checks settle that in about a millisecond, so moderation adds no noticeable wait to posting. The model only reads what they leave open. What lands in review waits in a queue, in your dashboard or through the API, and a signed webhook tells your site when somebody approves it.
How to detect harassment and pile-ons
Harassment is two different problems. One person insulting another is visible in one message, and ToxicFilter scores it under harassment when the insult is aimed at the reader rather than at a situation ("this game is shit" is not "you are an idiot"). A pile-on is not visible in any one message: it is many different people doing the same thing to one person. Send the thread to /v1/conversation with an author on each message and the number of distinct hostile people becomes part of the verdict, with a reason that says how many there were.
Disguised insults: why a word list is not enough
A list of bad words catches the people who were not trying. Everybody else writes f*ck, f u c k, fυck with a Greek upsilon or a zero-width space in the middle. ToxicFilter folds all of that back into plain letters before searching, and reports how much of a message was disguised as its own finding, under evasion. An allowlist blanks words out before anything searches them, so Scunthorpe, cockpit and classic are never caught.
Threats, hate and violence
A threat aimed at the reader ("I know where you live") is refused under violence. Attacks on groups land under hate, mostly through the model: the open word lists deliberately carry no slurs, because a complete enumeration of them has exactly one other use, and you can point ToxicFilter at your own list. A good record on your platform never lowers the line for either.
Self-harm and suicide: hold, never delete
A post about wanting to die is somebody who may be asking for help, and removing it removes them. self_harm holds for review and never blocks unless you decide otherwise, the answer says why in words, and the moderation.review webhook can alert whoever on your side is trained to respond. It is the one category where the right default is the gentle one.
Start from the community template, which holds profanity for a look sooner and still never refuses it for that alone, and refuses harassment, hate, threats and approaches to children sooner than the shipped lines do. Then try any change as a second policy running in parallel with the one in force: both verdicts are computed on your own traffic, only one acts, and the dashboard shows where they disagree before you switch.