Call the API before a comment appears under the article, and act on one of three answers: publish it, hold it for a person, or refuse it. Most comments are a reader saying something ordinary about the story, and the free checks settle those in about a millisecond, so moderation adds no wait that a reader would notice. The model only reads what they leave open, and you decide per call whether it may. A blocked comment comes back with the reason in a sentence your moderators can show the reader.
Threats against journalists and people named in a story
A threat aimed at the person reading ("I know where you live", "you better watch your back", "I will find you") is refused under violence by the free checks, in eight languages, and the community template refuses it sooner than the shipped line. The free lists hold the known phrasings; a threat worded some other way is the model's to read, with effort: high. A good record on your site never lowers the line for threats, so a regular commenter does not earn slack on one.
Attacks on groups land under hate, and that is mostly the model's job. The open word lists deliberately carry no slurs, so reading an attack on a group is the model's work. If your comment section draws that kind of traffic, let the model read the comments on the stories that attract it, or load your own list from outside the code. The community template refuses hate at 0.55, sooner than the shipped 0.70.
Heated argument that must stay up
People come to a comment section to disagree, often loudly. ToxicFilter separates what is said about a public figure, a policy or a story from what is aimed at the reader: "the mayor is a disaster" passes, "you are an idiot" is harassment. Swearing on its own is reported as toxicity, which the community template holds for a look and never refuses for that alone, so an angry reader who swears is reviewed rather than silenced.
A comment that is mostly links, or uses link shorteners, is scored under spam and held. A campaign is a different signal, and it is not in any one comment: it is the same comment arriving again and again. ToxicFilter keeps count across your own traffic, per project and never across other customers, of comments over about forty characters, recognises a copy that has been rewritten, and lets that count move the verdict on the next one. A repeat answered from the cache costs the check alone, one credit, whatever the first copy cost.
The review queue and statements of reasons
What lands in review waits in a queue, in your dashboard or through the API, and the moderation.resolved webhook tells your site when a moderator decides. Every blocked comment can carry the statement of reasons article 17 of the Digital Services Act asks for, in the reader's language, with the appeal behind it.