Gaming

Keep the trash talk, lose the abuse

Match chat, lobbies, guild channels and community servers, checked in about a millisecond as players type. ToxicFilter reads past the zeros and asterisks, sees a whole team turning on one player, holds the gold seller's shortened link and recognises an adult working on a child, while "gg, they destroyed us" goes straight through.

What goes wrong on platforms like yours

Chat that cannot wait

A message in a match is read within a second of being sent. A moderation call that takes longer than that is a chat nobody uses, so most teams end up filtering after the fact, when the player it was aimed at has already read it and left the game.

Insults spelled to get past the filter

"y0u are a f*cking 1d10t", a Cyrillic letter in the middle of a word, a space between every letter. Players learn what a word list looks for within one evening, and a filter that matches raw text only catches the ones who were not trying.

A whole team on one player

Four teammates each type one rude line at the player who missed the shot. Read one at a time, every message is ordinary frustration; together they are a pile-on, and it is the reason that player does not queue again.

Gold sellers, cheat shops and grown-ups who should not be there

Bots post cheap gold, boosted accounts and cheats behind shortened links in every busy channel. And on platforms full of children, an adult moving a chat to another app with gifts and secrecy is not one bad message: it is a sequence that has to reach a person.

What it decides, on content like yours

Judged right now by the engine the API uses, under the template for this kind of platform, with no model involved.

  1. 01 Trash talk after a match
  2. 02 An insult in disguise
  3. 03 A team turning on one player
  4. 04 A gold seller's link
  5. 05 An adult and a child in chat

Trash talk after a match POST /v1/text

gg, they destroyed us. Next round I'm killing you all, lol

allow 2 ms

Losing a round and promising revenge is how players talk. No insult, no threat phrase and nothing aimed at a person, so the free checks let it through in about a millisecond.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 2
}

An insult in disguise POST /v1/text

y0u are a f*cking 1d10t, uninstall

block 2 ms
  • Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.
  • Contains 1 insult(s), aimed at the reader.

The zeros, the ones and the asterisk are folded back into letters before anything is matched, so the disguise changes nothing. The insult aimed at the reader is what refuses it; the swear word on its own would only be held.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment",
    "toxicity"
  ],
  "scores": {
    "harassment": 0.8,
    "toxicity": 0.55
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.55,
      "reason": "Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "f*cking"
      ]
    },
    {
      "category": "harassment",
      "score": 0.8,
      "reason": "Contains 1 insult(s), aimed at the reader.",
      "evidence": [
        "idiot"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

A team turning on one player POST /v1/conversation

ana gg, that was close

bob you are a useless idiot, uninstall

cris lol you suck, idiot

dan you are such a noob, idiot

eve seriously, just leave the match

review 18 ms
  • 3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.

Three different people go after the same player in five messages. The last line alone passes; the thread is what holds it for a person, because one angry teammate is an argument and three are a pile-on.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.6
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.6,
      "reason": "3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 18
}

A gold seller's link POST /v1/text

Cheap gold, rare skins and boosted accounts, instant delivery: https://bit.ly/cheap-gold

review 7 ms
  • Uses a link shortener, which hides where the link goes.

What holds it is the shortened link, which hides where it goes; the same message posted again and again from one project is counted on its own, as repetition.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "spam"
  ],
  "scores": {
    "spam": 0.7
  },
  "signals": [
    {
      "category": "spam",
      "score": 0.7,
      "reason": "Uses a link shortener, which hides where the link goes.",
      "evidence": [
        "https://bit.ly/cheap-gold"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 7
}

An adult and a child in chat POST /v1/conversation

kid gg! i'm 12 btw, my mum lets me play till 9

max nice, you're good for 12. add me on snapchat, I'll gift you skins

max don't tell your parents, it's our secret

block 6 ms
  • Somebody in this conversation has said they are 12. The other participant asked for it to be kept secret, asked to move to another app. This needs a person now.

The child saying they are twelve is never a finding on its own. Another player asking for secrecy and a move to another app, after that, is the shape of an approach, and it is refused for a person to look at now.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "minor_safety"
  ],
  "scores": {
    "minor_safety": 0.9
  },
  "signals": [
    {
      "category": "minor_safety",
      "score": 0.9,
      "reason": "Somebody in this conversation has said they are 12. The other participant asked for it to be kept secret, asked to move to another app. This needs a person now.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 6
}

What it decides, on content like yours

Judged right now by the engine the API uses, under the template for this kind of platform, with no model involved.

Trash talk after a match POST /v1/text

gg, they destroyed us. Next round I'm killing you all, lol

allow 2 ms

Losing a round and promising revenge is how players talk. No insult, no threat phrase and nothing aimed at a person, so the free checks let it through in about a millisecond.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 2
}

An insult in disguise POST /v1/text

y0u are a f*cking 1d10t, uninstall

block 2 ms
  • Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.
  • Contains 1 insult(s), aimed at the reader.

The zeros, the ones and the asterisk are folded back into letters before anything is matched, so the disguise changes nothing. The insult aimed at the reader is what refuses it; the swear word on its own would only be held.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment",
    "toxicity"
  ],
  "scores": {
    "harassment": 0.8,
    "toxicity": 0.55
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.55,
      "reason": "Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "f*cking"
      ]
    },
    {
      "category": "harassment",
      "score": 0.8,
      "reason": "Contains 1 insult(s), aimed at the reader.",
      "evidence": [
        "idiot"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

A team turning on one player POST /v1/conversation

ana gg, that was close

bob you are a useless idiot, uninstall

cris lol you suck, idiot

dan you are such a noob, idiot

eve seriously, just leave the match

review 18 ms
  • 3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.

Three different people go after the same player in five messages. The last line alone passes; the thread is what holds it for a person, because one angry teammate is an argument and three are a pile-on.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.6
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.6,
      "reason": "3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 18
}

A gold seller's link POST /v1/text

Cheap gold, rare skins and boosted accounts, instant delivery: https://bit.ly/cheap-gold

review 7 ms
  • Uses a link shortener, which hides where the link goes.

What holds it is the shortened link, which hides where it goes; the same message posted again and again from one project is counted on its own, as repetition.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "spam"
  ],
  "scores": {
    "spam": 0.7
  },
  "signals": [
    {
      "category": "spam",
      "score": 0.7,
      "reason": "Uses a link shortener, which hides where the link goes.",
      "evidence": [
        "https://bit.ly/cheap-gold"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 7
}

An adult and a child in chat POST /v1/conversation

kid gg! i'm 12 btw, my mum lets me play till 9

max nice, you're good for 12. add me on snapchat, I'll gift you skins

max don't tell your parents, it's our secret

block 6 ms
  • Somebody in this conversation has said they are 12. The other participant asked for it to be kept secret, asked to move to another app. This needs a person now.

The child saying they are twelve is never a finding on its own. Another player asking for secrecy and a move to another app, after that, is the shape of an approach, and it is refused for a person to look at now.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "minor_safety"
  ],
  "scores": {
    "minor_safety": 0.9
  },
  "signals": [
    {
      "category": "minor_safety",
      "score": 0.9,
      "reason": "Somebody in this conversation has said they are 12. The other participant asked for it to be kept secret, asked to move to another app. This needs a person now.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 6
}

What it looks for

Each a score of its own, with its own lines, and the reason in a sentence whenever one acts.

The template you start from

Pick it when you create a policy and these rules are written for you, ready to edit. Everything it does not mention keeps following our defaults.

  • Toxicity holds at 0.30 · never refuses
  • Harassment holds at 0.30 · refuses at 0.65
  • Hate speech holds at 0.25 · refuses at 0.55
  • Violence holds at 0.25 · refuses at 0.60
  • Approach to a minor holds at 0.20 · refuses at 0.70
  • Filter evasion holds at 0.50 · refuses at 0.85
  • Spam holds at 0.50 · refuses at 0.80

Set by the template

One call before you publish

Send the text with where it will appear, and act on the decision.

The request

curl https://toxicfilter.com/api/v1/text \
  -H "Authorization: Bearer $TOXICFILTER_KEY" \
  -d content="gg, they destroyed us. Next round I'm killing you all, lol" \
  -d surface=comment
The full reference →

A month, in numbers

Items checked
600,000
Read by the model
15,000
Images
3,000
Credits, about
735,000

Fits in the Max plan. See the plans →

A guide to moderating game chat

What to check, how fast it has to be and where to set the lines, for match chat, lobbies and community servers.

How to moderate in-game chat in real time

Call the API when a player sends a message and act on one of three answers before it reaches the others: show it, hold it, or refuse it. The free checks settle a short chat line in about a millisecond, which is nearly every message in a match, so moderation adds no wait anybody notices. For the busiest channels, keep the model out with effort: low; elsewhere the default only has it read what the free checks leave in doubt, and you decide per call. What lands in review waits in a queue, in your dashboard or through the API, and a signed webhook tells your game when somebody decides.

Trash talk or harassment: where the line is

Game chat is full of words that look violent and are not. "We got destroyed", "I'm killing you all next round" and "gg" are gameplay. ToxicFilter acts on an insult aimed at the reader under harassment ("you idiot" is not "this map is stupid"), on a recognised threat under violence ("I know where you live", "kys"), and on attacks on groups under hate. A swear word on its own is scored under toxicity, which never blocks. The community template holds it for a look; a game chat usually wants that line higher, so a "that clutch was brilliant" with a swear word in it does not wait for anybody.

Disguised insults and why a word list is not enough

Players write f*ck, 1d10t, a Greek letter in place of a Latin one or a zero-width space in the middle of a word. ToxicFilter folds all of that back into plain letters before searching, and reports how much of the message was disguised as its own finding, under evasion. An allowlist blanks words out before anything searches them, so an innocent word that contains a bad one is never caught. Slurs are not shipped in the open lists; you load your own.

Pile-ons in match chat

A pile-on is invisible in any single message. Send the recent chat to /v1/conversation with a player id on each line, and ToxicFilter counts how many different players are being hostile to somebody in the second person. One furious teammate writing ten lines is one person; three or more different people is a pile-on, and the reason says how many there were. Nothing about the thread is stored, only the verdict on the last line.

Children on game platforms

The detector looks for the shape of an approach across a conversation: secrecy, a move to another app, photos, gifts, isolation, and another participant's stated age. A child saying "I'm 12" is never used against them, and a stated age alone is never a finding. With a stated minor, an approach is refused under minor_safety and should reach a person immediately. The instant checks read these shapes in eight languages, and with effort: high the model reads the whole conversation for what the wording leaves to context.

Choosing the thresholds for your game

Start from the community template: it refuses harassment, hate, threats and approaches to children sooner than the shipped lines. Then move what your players need, most often the toxicity line, and try the change as a second policy running beside the one in force. Both verdicts are computed on your own chat, only one acts, and the dashboard shows where they disagree before you switch.

Frequently asked questions

Is it fast enough for real-time game chat?

About a millisecond for a short chat message when the free checks settle it, measured as the whole request, which is nearly all chat traffic. The model only reads what they leave open, and you decide per call whether it may, so for a match chat you can keep it off and send only the uncertain ones to a person or to a slower second pass.

Will it remove normal competitive trash talk?

Not by itself. "gg, they destroyed us" or "next round I'm killing you all" carry no insult aimed at a person and no threat phrase, and pass. What gets acted on is an insult pointed at the reader, a recognised threat such as "I know where you live", or several people turning on one player. The community template does hold swearing for a look; for a game chat you will probably want to raise that line.

Can players get around it with leetspeak or symbols?

Much less than with a word list. Leetspeak, asterisks used as censor bars, look-alike letters from other alphabets, fullwidth characters, spaced-out letters and invisible characters are all folded back before anything is matched, and how much of a message was disguised is reported on its own, as evasion.

Does it catch slurs?

The word lists in the open repository deliberately carry no slurs, because a complete list of them has exactly one other use. You can point ToxicFilter at your own list, kept outside version control, and with the model reading, attacks on groups land under hate, which the community template refuses sooner than the shipped line.

How does it protect children on a game platform?

It reads the conversation for the shape of an approach: asking for secrecy, moving to another app, asking for photos, offering gifts or money, isolating somebody, beside another participant's stated age. A child saying their own age is never a finding. A match beside a stated minor is refused and should reach a person immediately. The patterns cover eight languages, and with effort high the model reads the whole conversation for what the wording leaves to context.

What about bots selling gold, accounts or cheats?

Their shape: shortened links, a message that is mostly links, and the same text, or the same text lightly rewritten, arriving many times in one project, which is counted across calls. The words your sellers use (gold, boosting, the cheat's name) go in a rule of your own words, and the model reads the offer itself when effort high brings it in.

Try it on your own traffic

2,000 credits a month on the free plan, no card. Enough to send a week of your own content and see what it says about it.