How to moderate in-game chat in real time
Call the API when a player sends a message and act on one of three answers before it reaches the others: show it, hold it, or refuse it. The free checks settle a short chat line in about a millisecond, which is nearly every message in a match, so moderation adds no wait anybody notices. For the busiest channels, keep the model out with effort: low; elsewhere the default only has it read what the free checks leave in doubt, and you decide per call. What lands in review waits in a queue, in your dashboard or through the API, and a signed webhook tells your game when somebody decides.
Trash talk or harassment: where the line is
Game chat is full of words that look violent and are not. "We got destroyed", "I'm killing you all next round" and "gg" are gameplay. ToxicFilter acts on an insult aimed at the reader under harassment ("you idiot" is not "this map is stupid"), on a recognised threat under violence ("I know where you live", "kys"), and on attacks on groups under hate. A swear word on its own is scored under toxicity, which never blocks. The community template holds it for a look; a game chat usually wants that line higher, so a "that clutch was brilliant" with a swear word in it does not wait for anybody.
Disguised insults and why a word list is not enough
Players write f*ck, 1d10t, a Greek letter in place of a Latin one or a zero-width space in the middle of a word. ToxicFilter folds all of that back into plain letters before searching, and reports how much of the message was disguised as its own finding, under evasion. An allowlist blanks words out before anything searches them, so an innocent word that contains a bad one is never caught. Slurs are not shipped in the open lists; you load your own.
Pile-ons in match chat
A pile-on is invisible in any single message. Send the recent chat to /v1/conversation with a player id on each line, and ToxicFilter counts how many different players are being hostile to somebody in the second person. One furious teammate writing ten lines is one person; three or more different people is a pile-on, and the reason says how many there were. Nothing about the thread is stored, only the verdict on the last line.
The detector looks for the shape of an approach across a conversation: secrecy, a move to another app, photos, gifts, isolation, and another participant's stated age. A child saying "I'm 12" is never used against them, and a stated age alone is never a finding. With a stated minor, an approach is refused under minor_safety and should reach a person immediately. The instant checks read these shapes in eight languages, and with effort: high the model reads the whole conversation for what the wording leaves to context.
Choosing the thresholds for your game
Start from the community template: it refuses harassment, hate, threats and approaches to children sooner than the shipped lines. Then move what your players need, most often the toxicity line, and try the change as a second policy running beside the one in force. Both verdicts are computed on your own chat, only one acts, and the dashboard shows where they disagree before you switch.