Multilingual moderation

One call for every language your users write in

The instant checks carry word lists and phrase patterns in eight languages, and the model reads far more. Tell ToxicFilter which languages your site speaks and it also notices the wall of an alphabet nobody there reads. And for a language beyond the eight, the model reads it when you ask, and the answer says whether it did.

How it works

  1. Say what your site speaks

    Send `locales` with the call, for example `["en"]`. It does not pick the word lists, which are all searched every time. It is what lets two checks tell a message that belongs from one that does not.

  2. The text is folded first

    Censor stars, spaced letters, leetspeak, missing accents and look-alike letters from another alphabet are undone before anything is searched, word by word, so genuine Russian or Greek is never rewritten into something it is not.

  3. Eight languages, in about a millisecond

    Insults, threats, slurs, scam shapes and personal data are matched in English, Spanish, Portuguese, French, Italian, German, Catalan and Dutch, whichever of them the message is in.

  4. The model for the rest

    With `effort: high` the model reads the message too, in far more languages than the lists cover, and the answer says whether it did.

See it decide

  1. 01 A Portuguese insult on an English site
  2. 02 A threat in German
  3. 03 A wall of Cyrillic spam
  4. 04 An insult spelled with Cyrillic letters
  5. 05 A short thank-you in Italian

A Portuguese insult on an English site POST /v1/text

Você é um idiota, ninguém te aguenta.

block 3 ms
  • Contains 1 insult(s), aimed at the reader.

The word lists are not chosen by `locales`: all eight are searched on every call, so an insult in Portuguese is caught on an English forum just as it would be on a Brazilian one.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.8
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.8,
      "reason": "Contains 1 insult(s), aimed at the reader.",
      "evidence": [
        "idiota"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 3
}

A threat in German POST /v1/text

Ich weiß, wo du wohnst. Ich bring dich um.

block 2 ms
  • Contains 2 phrase(s) threatening harm, aimed at the reader.

A threat aimed at the reader is refused under the community template whatever language it arrives in, as long as it is one of the eight.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "violence"
  ],
  "scores": {
    "violence": 0.95
  },
  "signals": [
    {
      "category": "violence",
      "score": 0.95,
      "reason": "Contains 2 phrase(s) threatening harm, aimed at the reader.",
      "evidence": [
        "ich weiss wo du wohnst",
        "bring dich um"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

A wall of Cyrillic spam POST /v1/text

Лучшие казино онлайн, бонус без депозита, заходи прямо сейчас https://example.com

block 11 ms
  • Written mostly in cyrillic (51 letters, 100% of the text) on a site set to en.
  • Reads as ru on a site set to en, and by a clear margin (0.17).

None of these words is on a list. What gives it away is that it is entirely Cyrillic on a site set to English. Without `locales` it is left alone, because a multilingual site posting Russian is doing its job.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "spam",
    "off_topic"
  ],
  "scores": {
    "spam": 0.97,
    "off_topic": 0.828
  },
  "signals": [
    {
      "category": "spam",
      "score": 0.97,
      "reason": "Written mostly in cyrillic (51 letters, 100% of the text) on a site set to en.",
      "evidence": [
        "Лучшие казино онлайн"
      ]
    },
    {
      "category": "off_topic",
      "score": 0.828,
      "reason": "Reads as ru on a site set to en, and by a clear margin (0.17).",
      "evidence": [
        "Лучшие казино онлайн, бонус без депозита, заходи прямо сейчас https://example.com"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 11
}

An insult spelled with Cyrillic letters POST /v1/text

You are an іdіоt and everyone knows it.

block 2 ms
  • Contains 1 insult(s), aimed at the reader.
  • About 16% of the letters are not the letters they appear to be.

Three of the letters in the insult are Cyrillic twins of Latin ones. A word that mixes alphabets is unmasked, and the answer adds that some letters are not what they appear to be.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.8,
    "evasion": 0.221
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.8,
      "reason": "Contains 1 insult(s), aimed at the reader.",
      "evidence": [
        "idiot"
      ]
    },
    {
      "category": "evasion",
      "score": 0.221,
      "reason": "About 16% of the letters are not the letters they appear to be.",
      "evidence": [
        "you are an idiot and everyone knows it"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

A short thank-you in Italian POST /v1/text

Grazie, ottimo articolo! Lo condivido con i miei colleghi.

allow 1 ms

Under sixty characters the language check does not run at all, because a guess about so few words is worse than no guess. A reader writing in their own language is not a finding.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 1
}

See it decide

A Portuguese insult on an English site POST /v1/text

Você é um idiota, ninguém te aguenta.

block 3 ms
  • Contains 1 insult(s), aimed at the reader.

The word lists are not chosen by `locales`: all eight are searched on every call, so an insult in Portuguese is caught on an English forum just as it would be on a Brazilian one.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.8
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.8,
      "reason": "Contains 1 insult(s), aimed at the reader.",
      "evidence": [
        "idiota"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 3
}

A threat in German POST /v1/text

Ich weiß, wo du wohnst. Ich bring dich um.

block 2 ms
  • Contains 2 phrase(s) threatening harm, aimed at the reader.

A threat aimed at the reader is refused under the community template whatever language it arrives in, as long as it is one of the eight.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "violence"
  ],
  "scores": {
    "violence": 0.95
  },
  "signals": [
    {
      "category": "violence",
      "score": 0.95,
      "reason": "Contains 2 phrase(s) threatening harm, aimed at the reader.",
      "evidence": [
        "ich weiss wo du wohnst",
        "bring dich um"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

A wall of Cyrillic spam POST /v1/text

Лучшие казино онлайн, бонус без депозита, заходи прямо сейчас https://example.com

block 11 ms
  • Written mostly in cyrillic (51 letters, 100% of the text) on a site set to en.
  • Reads as ru on a site set to en, and by a clear margin (0.17).

None of these words is on a list. What gives it away is that it is entirely Cyrillic on a site set to English. Without `locales` it is left alone, because a multilingual site posting Russian is doing its job.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "spam",
    "off_topic"
  ],
  "scores": {
    "spam": 0.97,
    "off_topic": 0.828
  },
  "signals": [
    {
      "category": "spam",
      "score": 0.97,
      "reason": "Written mostly in cyrillic (51 letters, 100% of the text) on a site set to en.",
      "evidence": [
        "Лучшие казино онлайн"
      ]
    },
    {
      "category": "off_topic",
      "score": 0.828,
      "reason": "Reads as ru on a site set to en, and by a clear margin (0.17).",
      "evidence": [
        "Лучшие казино онлайн, бонус без депозита, заходи прямо сейчас https://example.com"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 11
}

An insult spelled with Cyrillic letters POST /v1/text

You are an іdіоt and everyone knows it.

block 2 ms
  • Contains 1 insult(s), aimed at the reader.
  • About 16% of the letters are not the letters they appear to be.

Three of the letters in the insult are Cyrillic twins of Latin ones. A word that mixes alphabets is unmasked, and the answer adds that some letters are not what they appear to be.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.8,
    "evasion": 0.221
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.8,
      "reason": "Contains 1 insult(s), aimed at the reader.",
      "evidence": [
        "idiot"
      ]
    },
    {
      "category": "evasion",
      "score": 0.221,
      "reason": "About 16% of the letters are not the letters they appear to be.",
      "evidence": [
        "you are an idiot and everyone knows it"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

A short thank-you in Italian POST /v1/text

Grazie, ottimo articolo! Lo condivido con i miei colleghi.

allow 1 ms

Under sixty characters the language check does not run at all, because a guess about so few words is worse than no guess. A reader writing in their own language is not a finding.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 1
}

Moderating content in several languages, explained

What the instant checks read, what the model reads, and what a clean answer means in each case.

How to moderate comments in several languages

Send every comment to the same endpoint, whatever language it is in, with locales set to the languages your site speaks. The word lists for English, Spanish, Portuguese, French, Italian, German, Catalan and Dutch are all searched on every call, so you do not have to detect the language first or route a message to a different model. A short comment the free checks settle comes back in about a millisecond; the model reads it as well when you send "effort": "high".

Which languages the instant checks carry

The vocabulary behind toxicity, harassment, hate, sexual content, violence, self-harm and scams is written in eight languages, and so are the phrase patterns for pile-ons, approaches to children, deals taken off the platform and the scam shapes. Every group must have every language, and a test fails the build when one is missing, because a half-added language looks covered when it is not. The checks that do not depend on words (links, wallet addresses, card and IBAN checksums, repeated messages) work in any language.

Spam in another alphabet

The commonest backlink attack is a wall of Cyrillic dropped into the comments of a site where nobody reads it. With locales set, a text in an alphabet those languages do not use is reported as spam when it has at least fifteen letters of that alphabet and more than one per cent of the text. Links are removed before counting, because a web address is not written in any language and counted as Latin it hid a short Russian bio. A quoted Russian word clears neither gate. Without locales the check stays silent: a multilingual site posting Russian is doing its job.

A message in the wrong language

Fluent English spam on a Spanish forum uses the right alphabet and the wrong language. Over sixty characters, the language check compares the language the text reads as against the ones you declared and reports off_topic only when the winner beats yours by a clear margin, so a Spanish post full of English game titles is left alone. It holds for review and never refuses, since a visitor writing in their own language has done nothing wrong.

Disguised words and genuine foreign script

Nothing is matched against raw text. Look-alike letters are folded word by word, and only where a word mixes alphabets, or sits in a Latin text made of foreign letters that all have a Latin twin, so іdіоt with Cyrillic letters is unmasked while a sentence in Russian keeps its letters. The fold never transliterates to ASCII either, which would erase Chinese, Arabic, Hebrew, Japanese and Korean entirely and make every text in those scripts look like gibberish.

Languages beyond the eight

Outside these eight languages the structural checks still run on every call, and the model reads the language itself. The languages page of the docs sets out what each layer reads. On a site in a language that is not listed, send "effort": "high" for anything that matters, and add the words that matter to you to your own policy lists.

How the work is split

The instant checks settle the clear cases in about a millisecond, the model reads what depends on context, and your rules and your people have the last word.

  • Eight languages on every call

    The word lists and phrase patterns for all eight languages are searched on every call, whatever the site speaks, and the structural checks (links, wallets, card and IBAN checksums, repetition) work in any language.

  • Context is the model's job

    A sentence that depends on its context, a phrasing nobody has written down or a language beyond the eight is read by the model, which effort high brings in. The answer says whether it actually read.

  • Your words, in any language

    Your policy's block, review and allow lists work in any language and go through the same folding, so a community in any language teaches the instant checks its own vocabulary.

Frequently asked questions

Which languages does ToxicFilter moderate?

The instant checks carry word lists and phrase patterns in English, Spanish, Portuguese, French, Italian, German, Catalan and Dutch. The structural checks (links, wallet addresses, card and IBAN checksums, repetition) work in any language. The model, when you ask for it, reads far more languages than the lists cover.

Do I have to tell it which language a comment is in?

No. All eight word lists are searched on every call, because a comment in Portuguese arrives on a Spanish forum every day. `locales` tells it which languages your site speaks, which is what makes the alphabet check and the language check possible.

What happens to a comment in a language my site does not use?

Over sixty characters, a message that reads as another language by a clear margin is reported as `off_topic`, which holds it for review and never refuses it. A long block in an alphabet your languages do not use (Cyrillic, Greek, Chinese, Arabic or Hebrew) is reported as spam: almost entirely foreign it is refused, a smaller share is held.

Does it catch insults written with letters from another alphabet?

Yes. A word that mixes alphabets, like Latin letters with Cyrillic twins in the middle, is folded back before the search and reported as disguised. A word written entirely in a real foreign alphabet is left as it is, which is why genuine Russian is never flagged as an attack.

Can I add words in my own language?

Yes. Your policy's block, review and allow lists work in any language and go through the same folding, so you write a word once and its spaced, starred and accentless spellings are found too. For a community whose language is not among the eight, that is how the free path learns your words.

Is the model needed for languages outside the eight?

For anything that matters, yes. There the structural checks and your own word lists run on the instant path, and the model reads the language itself. With `effort: high` the model reads the text, and the answer says whether it actually did.

Try it on your own traffic

2,000 credits a month on the free plan, no card. Enough to send a week of your own content and see what it says about it.