News publishers

Keep the comments open, and the threats out of them

Every comment under every article, checked before it appears. ToxicFilter refuses the threat aimed at the reporter and the insult aimed at another reader, holds the link campaign and the pile-on for a person, and leaves the angry, legitimate argument about the story where it is.

What goes wrong on platforms like yours

Threats against the newsroom

A reporter writes about a local contract and the comments fill with "I know where you live". Threats against journalists, and against the people named in a story, arrive under the article itself, in public, and every hour they stay up is an hour the target reads them.

Readers turning on each other

A comment section is a crowd. One reader insults another, three more join in, and the person who disagreed stops commenting. Read one at a time, the replies are bad temper; together they decide who gets to speak under your articles.

Link campaigns in the busiest stories

The story that is on every front page is the one the spammers find first: "great article, more here" with three shortened links, then the same comment, slightly rewritten, under every other piece that day.

Heated argument is the point

Readers are angry about the mayor, the budget and the war, and they say so. A filter that cannot tell "this plan is a disgrace" from an attack on a person either removes the debate your readers came for or lets the attacks through to keep it.

What it decides, on content like yours

Judged right now by the engine the API uses, under the template for this kind of platform, with no model involved.

  1. 01 A threat at the reporter
  2. 02 An insult at another reader
  3. 03 A pile-on under the article
  4. 04 A link campaign
  5. 05 Angry, and legitimate

A threat at the reporter POST /v1/text

Nice piece, reporter. I know where you live. Watch your back.

block 10 ms
  • Contains 1 phrase(s) threatening harm, aimed at the reader.

Two threatening phrases aimed at the person reading, refused under violence by the free checks with no model involved. The praise in front of them changes nothing.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "violence"
  ],
  "scores": {
    "violence": 0.95
  },
  "signals": [
    {
      "category": "violence",
      "score": 0.95,
      "reason": "Contains 1 phrase(s) threatening harm, aimed at the reader.",
      "evidence": [
        "i know where you live"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 10
}

An insult at another reader POST /v1/text

You are a f*cking 1d10t, you know nothing about this town.

block 2 ms
  • Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.
  • Contains 1 insult(s), aimed at the reader.

An insult aimed at the reader, disguised with numbers and an asterisk that are folded back into letters before anything is matched. The community template refuses harassment sooner than the shipped line.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment",
    "toxicity"
  ],
  "scores": {
    "harassment": 0.8,
    "toxicity": 0.55
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.55,
      "reason": "Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "f*cking"
      ]
    },
    {
      "category": "harassment",
      "score": 0.8,
      "reason": "Contains 1 insult(s), aimed at the reader.",
      "evidence": [
        "idiot"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

A pile-on under the article POST /v1/conversation

ana The bridge has been shut for months and nobody explains why.

bob shut up, you idiot, you know nothing

cris lol you are such a loser

dan you are worthless, idiot

eve honestly, just leave

review 22 ms
  • 3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.

Three different readers on the one who asked a question. The last reply alone is a brusque sentence and passes; the thread is what holds it for a person.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.6
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.6,
      "reason": "3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 22
}

A link campaign POST /v1/text

Great article. More here: https://bit.ly/x1 https://bit.ly/x2 https://bit.ly/x3

review 7 ms
  • 3 links in about 16 words: mostly links, barely a message.
  • Uses a link shortener, which hides where the link goes.

Mostly links, all of them shortened, under a sentence of praise. Held rather than refused: one copy is not proof of a campaign, and the same comment arriving again from your traffic is what raises it.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "spam"
  ],
  "scores": {
    "spam": 0.75
  },
  "signals": [
    {
      "category": "spam",
      "score": 0.75,
      "reason": "3 links in about 16 words: mostly links, barely a message.",
      "evidence": [
        "https://bit.ly/x1",
        "https://bit.ly/x2",
        "https://bit.ly/x3"
      ]
    },
    {
      "category": "spam",
      "score": 0.7,
      "reason": "Uses a link shortener, which hides where the link goes.",
      "evidence": [
        "https://bit.ly/x1",
        "https://bit.ly/x2",
        "https://bit.ly/x3"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 7
}

Angry, and legitimate POST /v1/text

This mayor is a disaster and his housing plan is a joke. Vote him out in May.

allow 8 ms

Harsh about a public figure and his policy, aimed at nobody in the thread. It stays up, settled by the free checks in about a millisecond.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 8
}

What it decides, on content like yours

Judged right now by the engine the API uses, under the template for this kind of platform, with no model involved.

A threat at the reporter POST /v1/text

Nice piece, reporter. I know where you live. Watch your back.

block 10 ms
  • Contains 1 phrase(s) threatening harm, aimed at the reader.

Two threatening phrases aimed at the person reading, refused under violence by the free checks with no model involved. The praise in front of them changes nothing.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "violence"
  ],
  "scores": {
    "violence": 0.95
  },
  "signals": [
    {
      "category": "violence",
      "score": 0.95,
      "reason": "Contains 1 phrase(s) threatening harm, aimed at the reader.",
      "evidence": [
        "i know where you live"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 10
}

An insult at another reader POST /v1/text

You are a f*cking 1d10t, you know nothing about this town.

block 2 ms
  • Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.
  • Contains 1 insult(s), aimed at the reader.

An insult aimed at the reader, disguised with numbers and an asterisk that are folded back into letters before anything is matched. The community template refuses harassment sooner than the shipped line.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment",
    "toxicity"
  ],
  "scores": {
    "harassment": 0.8,
    "toxicity": 0.55
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.55,
      "reason": "Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "f*cking"
      ]
    },
    {
      "category": "harassment",
      "score": 0.8,
      "reason": "Contains 1 insult(s), aimed at the reader.",
      "evidence": [
        "idiot"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

A pile-on under the article POST /v1/conversation

ana The bridge has been shut for months and nobody explains why.

bob shut up, you idiot, you know nothing

cris lol you are such a loser

dan you are worthless, idiot

eve honestly, just leave

review 22 ms
  • 3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.

Three different readers on the one who asked a question. The last reply alone is a brusque sentence and passes; the thread is what holds it for a person.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.6
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.6,
      "reason": "3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 22
}

A link campaign POST /v1/text

Great article. More here: https://bit.ly/x1 https://bit.ly/x2 https://bit.ly/x3

review 7 ms
  • 3 links in about 16 words: mostly links, barely a message.
  • Uses a link shortener, which hides where the link goes.

Mostly links, all of them shortened, under a sentence of praise. Held rather than refused: one copy is not proof of a campaign, and the same comment arriving again from your traffic is what raises it.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "spam"
  ],
  "scores": {
    "spam": 0.75
  },
  "signals": [
    {
      "category": "spam",
      "score": 0.75,
      "reason": "3 links in about 16 words: mostly links, barely a message.",
      "evidence": [
        "https://bit.ly/x1",
        "https://bit.ly/x2",
        "https://bit.ly/x3"
      ]
    },
    {
      "category": "spam",
      "score": 0.7,
      "reason": "Uses a link shortener, which hides where the link goes.",
      "evidence": [
        "https://bit.ly/x1",
        "https://bit.ly/x2",
        "https://bit.ly/x3"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 7
}

Angry, and legitimate POST /v1/text

This mayor is a disaster and his housing plan is a joke. Vote him out in May.

allow 8 ms

Harsh about a public figure and his policy, aimed at nobody in the thread. It stays up, settled by the free checks in about a millisecond.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 8
}

What it looks for

Each a score of its own, with its own lines, and the reason in a sentence whenever one acts.

The template you start from

Pick it when you create a policy and these rules are written for you, ready to edit. Everything it does not mention keeps following our defaults.

  • Toxicity holds at 0.30 · never refuses
  • Harassment holds at 0.30 · refuses at 0.65
  • Hate speech holds at 0.25 · refuses at 0.55
  • Violence holds at 0.25 · refuses at 0.60
  • Approach to a minor holds at 0.20 · refuses at 0.70
  • Spam holds at 0.50 · refuses at 0.80
  • Filter evasion holds at 0.50 · refuses at 0.85

Set by the template

One call before you publish

Send the text with where it will appear, and act on the decision.

The request

curl https://toxicfilter.com/api/v1/text \
  -H "Authorization: Bearer $TOXICFILTER_KEY" \
  -d content="Nice piece, reporter. I know where you live. Watch your back." \
  -d surface=comment
The full reference →

A month, in numbers

Items checked
400,000
Read by the model
25,000
Images
2,000
Credits, about
595,000

Fits in the Max plan. See the plans →

A guide to moderating a comment section

What to check, when to check it and how to set the rules, for the comments under your articles.

How to moderate comments on a news site

Call the API before a comment appears under the article, and act on one of three answers: publish it, hold it for a person, or refuse it. Most comments are a reader saying something ordinary about the story, and the free checks settle those in about a millisecond, so moderation adds no wait that a reader would notice. The model only reads what they leave open, and you decide per call whether it may. A blocked comment comes back with the reason in a sentence your moderators can show the reader.

Threats against journalists and people named in a story

A threat aimed at the person reading ("I know where you live", "you better watch your back", "I will find you") is refused under violence by the free checks, in eight languages, and the community template refuses it sooner than the shipped line. The free lists hold the known phrasings; a threat worded some other way is the model's to read, with effort: high. A good record on your site never lowers the line for threats, so a regular commenter does not earn slack on one.

Hate speech in the comments

Attacks on groups land under hate, and that is mostly the model's job. The open word lists deliberately carry no slurs, so reading an attack on a group is the model's work. If your comment section draws that kind of traffic, let the model read the comments on the stories that attract it, or load your own list from outside the code. The community template refuses hate at 0.55, sooner than the shipped 0.70.

Heated argument that must stay up

People come to a comment section to disagree, often loudly. ToxicFilter separates what is said about a public figure, a policy or a story from what is aimed at the reader: "the mayor is a disaster" passes, "you are an idiot" is harassment. Swearing on its own is reported as toxicity, which the community template holds for a look and never refuses for that alone, so an angry reader who swears is reviewed rather than silenced.

A comment that is mostly links, or uses link shorteners, is scored under spam and held. A campaign is a different signal, and it is not in any one comment: it is the same comment arriving again and again. ToxicFilter keeps count across your own traffic, per project and never across other customers, of comments over about forty characters, recognises a copy that has been rewritten, and lets that count move the verdict on the next one. A repeat answered from the cache costs the check alone, one credit, whatever the first copy cost.

The review queue and statements of reasons

What lands in review waits in a queue, in your dashboard or through the API, and the moderation.resolved webhook tells your site when a moderator decides. Every blocked comment can carry the statement of reasons article 17 of the Digital Services Act asks for, in the reader's language, with the appeal behind it.

Frequently asked questions

Can it tell a heated political comment from abuse?

It looks at who the words are aimed at. Calling a policy a disgrace or a politician a liar is not aimed at the reader, and it passes. An insult or a threat aimed at the person reading ("you are an idiot", "I know where you live") is scored as harassment or violence. Swearing on its own is reported as toxicity: the community template holds it for a look and never refuses a comment for that alone.

Does it catch hate speech against groups?

Mostly through the model. The open word lists deliberately carry no slurs, because a complete list of them has exactly one other use, so an attack on a group is read by the model, which effort high brings in on every comment. You can load your own slur list from outside the code, and the community template refuses hate sooner than the shipped line.

How do you detect coordinated or repeated comments?

ToxicFilter counts, across your own traffic and never across other customers', how often the same comment of more than about forty characters has arrived in the last minutes, and recognises it rewritten with a link swapped or a word changed. Those counts move the verdict on the next copy. Under one article, send the thread to /v1/conversation with an author on each comment and it also counts how many different people are hostile to the same person.

Which languages does it cover?

The instant checks carry word and phrase lists in English, Spanish, Portuguese, French, Italian, German, Catalan and Dutch. The model reads many more. Outside those eight, effort high has the model read every comment.

Can our moderators approve a held comment?

Yes. Everything that lands in review waits in a queue, in your dashboard or through the API. When somebody approves or rejects it, a signed moderation.resolved webhook tells your site, so the comment goes live or stays down without anybody copying it across.

Do we have to tell a reader why their comment was removed?

In the EU, yes: article 17 of the Digital Services Act requires a statement of reasons for every removal or restriction, whatever the size of the publisher. ToxicFilter can write it with every blocked comment, in the reader's language, and take their appeal.

Try it on your own traffic

2,000 credits a month on the free plan, no card. Enough to send a week of your own content and see what it says about it.