Social networks

Let people talk, and stop the pile-on before it lands

Posts, comments, replies and direct messages, checked as they are written. ToxicFilter reads the thread as well as the message, so it sees four people turning on one, the insult spelled to slip past a word list and the giveaway that is a wallet address, and it sends somebody in crisis to a person instead of deleting them.

What goes wrong on platforms like yours

Pile-ons that no single message shows

Thirty accounts each write one rude sentence under somebody's post. Read one at a time, every reply is ordinary bad temper; together they are a pile-on, and the person on the other end of it leaves your platform. A filter that reads messages one by one never sees it happen.

Abuse spelled to get past the filter

"y0u are a f*cking 1d10t", a Cyrillic letter in the middle of a word, a zero-width space between two others. People who want to insult somebody learn in a day what a word list looks for, and a filter that matches raw text only catches the ones who were not trying.

Scam and spam campaigns in the replies

The crypto giveaway with a wallet address, the "free followers" with three shortened links, the same message rewritten a hundred times by the same network of accounts. They arrive in the busiest threads, because that is where the audience is.

People in crisis, and children

A post saying somebody wants to end their life is not a rule being broken, and deleting it is the worst possible answer. An adult moving a conversation with a thirteen-year-old to another app is not one bad message either. Both need a person, quickly, and neither is a word on a list.

What it decides, on content like yours

Judged right now by the engine the API uses, under the template for this kind of platform, with no model involved.

  1. 01 A pile-on
  2. 02 An insult in disguise
  3. 03 Somebody in crisis
  4. 04 A crypto giveaway
  5. 05 An ordinary comment

A pile-on POST /v1/conversation

bob you can't draw, you idiot

cris lol you are such a loser

dan you should quit, idiot

eve honestly, just stop posting

review 17 ms
  • 3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.

Three different people on the same person in four messages. The last reply alone is a rude sentence and passes; the thread is what holds it for a person, and with a fourth person the community template refuses it.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.638
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.638,
      "reason": "3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 17
}

An insult in disguise POST /v1/text

y0u are a f*cking 1d10t

block 1 ms
  • Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.
  • Contains 1 insult(s), aimed at the reader.
  • About 17% of the letters are not the letters they appear to be.

The zeros, the ones and the asterisk are folded back into letters before anything is matched, so the disguise changes nothing. How much of the text was disguised is reported on its own, as evasion.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment",
    "toxicity"
  ],
  "scores": {
    "harassment": 0.8,
    "toxicity": 0.55,
    "evasion": 0.243
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.55,
      "reason": "Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "f*cking"
      ]
    },
    {
      "category": "harassment",
      "score": 0.8,
      "reason": "Contains 1 insult(s), aimed at the reader.",
      "evidence": [
        "idiot"
      ]
    },
    {
      "category": "evasion",
      "score": 0.243,
      "reason": "About 17% of the letters are not the letters they appear to be.",
      "evidence": [
        "you are a f*cking idiot"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 1
}

Somebody in crisis POST /v1/text

I can't do this anymore. I want to end my life.

review 1 ms
  • Contains 1 phrase(s) about self-harm. This may be someone asking for help rather than breaking a rule.

Held for a person, never refused. The person harmed by deleting this post is the one who wrote it, which is why self-harm reviews by default and does not block.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "self_harm"
  ],
  "scores": {
    "self_harm": 0.7
  },
  "signals": [
    {
      "category": "self_harm",
      "score": 0.7,
      "reason": "Contains 1 phrase(s) about self-harm. This may be someone asking for help rather than breaking a rule.",
      "evidence": [
        "end my life"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 1
}

A crypto giveaway POST /v1/text

Check my profile for a free crypto giveaway 🚀🚀 send 0.1 BTC and get 1 BTC back! bc1qxy2kgdygjrsqtzq2n0yrf2493p83kkfjhx0wlh

block 8 ms
  • Contains a cryptocurrency wallet address, with an instruction to send to it or beside another scam shape.

A wallet address beside a promise to send more back is a scam whatever the emojis around it say. It is settled by the free checks, with no model involved.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "scam"
  ],
  "scores": {
    "scam": 0.85
  },
  "topics": {
    "crypto": 0.471
  },
  "signals": [
    {
      "category": "scam",
      "score": 0.85,
      "reason": "Contains a cryptocurrency wallet address, with an instruction to send to it or beside another scam shape.",
      "evidence": [
        "bc1qxy2kgdygjrsqtzq2n0yrf2493p83kkfjhx0wlh"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 8
}

An ordinary comment POST /v1/text

Great photo! Where was this taken? The light on the water is beautiful.

allow 6 ms

Nearly every comment looks like this. It is settled by the free checks in about a millisecond, and nothing is paid for.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 6
}

What it decides, on content like yours

Judged right now by the engine the API uses, under the template for this kind of platform, with no model involved.

A pile-on POST /v1/conversation

bob you can't draw, you idiot

cris lol you are such a loser

dan you should quit, idiot

eve honestly, just stop posting

review 17 ms
  • 3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.

Three different people on the same person in four messages. The last reply alone is a rude sentence and passes; the thread is what holds it for a person, and with a fourth person the community template refuses it.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.638
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.638,
      "reason": "3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 17
}

An insult in disguise POST /v1/text

y0u are a f*cking 1d10t

block 1 ms
  • Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.
  • Contains 1 insult(s), aimed at the reader.
  • About 17% of the letters are not the letters they appear to be.

The zeros, the ones and the asterisk are folded back into letters before anything is matched, so the disguise changes nothing. How much of the text was disguised is reported on its own, as evasion.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment",
    "toxicity"
  ],
  "scores": {
    "harassment": 0.8,
    "toxicity": 0.55,
    "evasion": 0.243
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.55,
      "reason": "Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "f*cking"
      ]
    },
    {
      "category": "harassment",
      "score": 0.8,
      "reason": "Contains 1 insult(s), aimed at the reader.",
      "evidence": [
        "idiot"
      ]
    },
    {
      "category": "evasion",
      "score": 0.243,
      "reason": "About 17% of the letters are not the letters they appear to be.",
      "evidence": [
        "you are a f*cking idiot"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 1
}

Somebody in crisis POST /v1/text

I can't do this anymore. I want to end my life.

review 1 ms
  • Contains 1 phrase(s) about self-harm. This may be someone asking for help rather than breaking a rule.

Held for a person, never refused. The person harmed by deleting this post is the one who wrote it, which is why self-harm reviews by default and does not block.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "self_harm"
  ],
  "scores": {
    "self_harm": 0.7
  },
  "signals": [
    {
      "category": "self_harm",
      "score": 0.7,
      "reason": "Contains 1 phrase(s) about self-harm. This may be someone asking for help rather than breaking a rule.",
      "evidence": [
        "end my life"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 1
}

A crypto giveaway POST /v1/text

Check my profile for a free crypto giveaway 🚀🚀 send 0.1 BTC and get 1 BTC back! bc1qxy2kgdygjrsqtzq2n0yrf2493p83kkfjhx0wlh

block 8 ms
  • Contains a cryptocurrency wallet address, with an instruction to send to it or beside another scam shape.

A wallet address beside a promise to send more back is a scam whatever the emojis around it say. It is settled by the free checks, with no model involved.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "scam"
  ],
  "scores": {
    "scam": 0.85
  },
  "topics": {
    "crypto": 0.471
  },
  "signals": [
    {
      "category": "scam",
      "score": 0.85,
      "reason": "Contains a cryptocurrency wallet address, with an instruction to send to it or beside another scam shape.",
      "evidence": [
        "bc1qxy2kgdygjrsqtzq2n0yrf2493p83kkfjhx0wlh"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 8
}

An ordinary comment POST /v1/text

Great photo! Where was this taken? The light on the water is beautiful.

allow 6 ms

Nearly every comment looks like this. It is settled by the free checks in about a millisecond, and nothing is paid for.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 6
}

What it looks for

Each a score of its own, with its own lines, and the reason in a sentence whenever one acts.

The template you start from

Pick it when you create a policy and these rules are written for you, ready to edit. Everything it does not mention keeps following our defaults.

  • Toxicity holds at 0.30 · never refuses
  • Harassment holds at 0.30 · refuses at 0.65
  • Hate speech holds at 0.25 · refuses at 0.55
  • Violence holds at 0.25 · refuses at 0.60
  • Approach to a minor holds at 0.20 · refuses at 0.70
  • Self-harm holds at 0.30 · never refuses
  • Filter evasion holds at 0.50 · refuses at 0.85
  • Scam holds at 0.45 · refuses at 0.75

Set by the template

One call before you publish

Send the text with where it will appear, and act on the decision.

The request

curl https://toxicfilter.com/api/v1/text \
  -H "Authorization: Bearer $TOXICFILTER_KEY" \
  -d content="you can't draw, you idiot" \
  -d surface=comment
The full reference →

A month, in numbers

Items checked
300,000
Read by the model
20,000
Images
10,000
Credits, about
540,000

Fits in the Max plan. See the plans →

A guide to moderating a social network

What to check, when to check it and how to set the rules, for posts, comments, replies and messages.

How to moderate comments and posts in real time

Call the API before a post, a comment or a message is published, and act on one of three answers: publish it, hold it for a person, or refuse it. Most of a community's traffic is someone saying something ordinary to someone else, and the free checks settle that in about a millisecond, so moderation adds no noticeable wait to posting. The model only reads what they leave open. What lands in review waits in a queue, in your dashboard or through the API, and a signed webhook tells your site when somebody approves it.

How to detect harassment and pile-ons

Harassment is two different problems. One person insulting another is visible in one message, and ToxicFilter scores it under harassment when the insult is aimed at the reader rather than at a situation ("this game is shit" is not "you are an idiot"). A pile-on is not visible in any one message: it is many different people doing the same thing to one person. Send the thread to /v1/conversation with an author on each message and the number of distinct hostile people becomes part of the verdict, with a reason that says how many there were.

Disguised insults: why a word list is not enough

A list of bad words catches the people who were not trying. Everybody else writes f*ck, f u c k, fυck with a Greek upsilon or a zero-width space in the middle. ToxicFilter folds all of that back into plain letters before searching, and reports how much of a message was disguised as its own finding, under evasion. An allowlist blanks words out before anything searches them, so Scunthorpe, cockpit and classic are never caught.

Threats, hate and violence

A threat aimed at the reader ("I know where you live") is refused under violence. Attacks on groups land under hate, mostly through the model: the open word lists deliberately carry no slurs, because a complete enumeration of them has exactly one other use, and you can point ToxicFilter at your own list. A good record on your platform never lowers the line for either.

Self-harm and suicide: hold, never delete

A post about wanting to die is somebody who may be asking for help, and removing it removes them. self_harm holds for review and never blocks unless you decide otherwise, the answer says why in words, and the moderation.review webhook can alert whoever on your side is trained to respond. It is the one category where the right default is the gentle one.

Choosing the thresholds for your community

Start from the community template, which holds profanity for a look sooner and still never refuses it for that alone, and refuses harassment, hate, threats and approaches to children sooner than the shipped lines do. Then try any change as a second policy running in parallel with the one in force: both verdicts are computed on your own traffic, only one acts, and the dashboard shows where they disagree before you switch.

Frequently asked questions

How do you detect a pile-on or coordinated harassment?

Send the thread to /v1/conversation, with an author id on each message. ToxicFilter counts how many different people in it are being hostile to somebody in the second person, and how much of the thread that is. One furious person writing fifteen messages is an argument and is not counted as a pile-on; four different people turning on the same person is, and the score says how many there were.

Can people get around it with leetspeak, symbols or look-alike letters?

Much less than with a word list. Everything is folded before it is matched: leetspeak, asterisks used as censor bars, Cyrillic and Greek letters that look Latin, fullwidth characters, accents, spaced-out letters and invisible characters. A word only counts as disguised when it mixes alphabets, so genuine Russian or Greek is left alone.

What happens to posts about self-harm or suicide?

They are held for review and never refused by default, because deleting them hurts the person who wrote them. The answer says the post may be somebody asking for help, and a webhook can alert whoever on your side can respond. A good record never lowers the attention a post like this gets.

Does it protect children on the platform?

It looks for the shape of an approach across a conversation: moving to another app, asking for secrecy, asking for photos, gifts, isolating somebody, and another participant's stated age. A child stating their own age is never a finding. A match beside a stated minor is refused and should reach a person immediately, and with effort high the model reads the whole conversation for what the wording leaves to context.

How fast is it on a busy feed?

About a millisecond for a short comment when the free checks settle it, which is nearly all traffic. The model only reads what they leave open, and you decide per call whether it may. Repeats, such as the same spam posted a hundred times, are answered from a cache.

Do I have to explain to a user why their post was removed?

In the EU, yes: article 17 of the Digital Services Act requires a statement of reasons for every removal or restriction, whatever the size of the platform. ToxicFilter can write it with every blocked post, in the user's language, and take their appeal.

Try it on your own traffic

2,000 credits a month on the free plan, no card. Enough to send a week of your own content and see what it says about it.