Text moderation

Every comment gets an answer, and the reason for it

Send a comment, a post or a message to /v1/text before you publish it. It comes back allow, review or block, scored across fifteen categories, with a sentence saying why and the fragment that made it say so. Most of it is settled in about a millisecond, without a model.

How it works

  1. You send the text

    One POST to /v1/text with the content, and optionally the languages you expect, where it is going to appear and the rules to apply. Your own id for it comes back with the answer.

  2. The free checks read it

    Word lists and phrase patterns in eight languages, run on a folded copy of the text, so f*ck, f u c k and a Greek letter in the middle of a word read as what they are. Links, contact details, flooding and the wrong alphabet are checked in the same pass.

  3. The model only if it adds something

    When the free checks settle it, nothing else runs and the check costs one credit. When they leave the question open, the model reads it and the tokens it used are added to the bill.

  4. You get a decision you can explain

    allow, review or block, a score per category, and signals that each name the category, the detector, a reason in words and the fragment that triggered it.

See it decide

  1. 01 An ordinary comment
  2. 02 A swear word in a compliment
  3. 03 An insult in disguise
  4. 04 Links and little else
  5. 05 A threat

An ordinary comment POST /v1/text

Lovely photos. Where is this lake? I would like to visit in spring.

allow 8 ms

Nearly every comment looks like this. The free checks settle it and nothing is paid beyond the check.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 8
}

A swear word in a compliment POST /v1/text

This trailer is fucking brilliant, I watched it five times.

allow 2 ms
  • Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.

The profanity is reported under toxicity, and on its own it only says the tone is casual. It is a score in one category, not a verdict on the whole comment.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "scores": {
    "toxicity": 0.35
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.35,
      "reason": "Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "fucking"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

An insult in disguise POST /v1/text

nobody asked you, you st*pid 1d10t. go back under your rock

block 2 ms
  • Contains 2 insult(s), aimed at the reader.

The asterisk and the digits are folded back before anything is matched, and the insult is aimed at the reader, so it is harassment and it is refused.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.85
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.85,
      "reason": "Contains 2 insult(s), aimed at the reader.",
      "evidence": [
        "you st*pid",
        "idiot"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

Links and little else POST /v1/text

Earn 300 dollars a day from home!!! Click now: https://bit.ly/3xCash and https://t.me/cashvip_now

review 8 ms
  • 2 links in about 18 words: mostly links, barely a message.
  • Uses a link shortener, which hides where the link goes.
  • Links to a chat invite or a free landing page, which is where promotion parks itself.

A link shortener, a chat invite and two links in eighteen words. Held for a person under the shipped lines; a rule in the call can make it a refusal.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "spam"
  ],
  "scores": {
    "spam": 0.7
  },
  "signals": [
    {
      "category": "spam",
      "score": 0.65,
      "reason": "2 links in about 18 words: mostly links, barely a message.",
      "evidence": [
        "https://bit.ly/3xCash",
        "https://t.me/cashvip_now"
      ]
    },
    {
      "category": "spam",
      "score": 0.7,
      "reason": "Uses a link shortener, which hides where the link goes.",
      "evidence": [
        "https://bit.ly/3xCash"
      ]
    },
    {
      "category": "spam",
      "score": 0.55,
      "reason": "Links to a chat invite or a free landing page, which is where promotion parks itself.",
      "evidence": [
        "https://t.me/cashvip_now"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 8
}

A threat POST /v1/text

I know where you live and I will find you. You will pay for this.

block 10 ms
  • Contains 2 phrase(s) threatening harm, aimed at the reader.

Two phrases threatening harm, aimed at the reader. Violence is one of the categories that refuses early.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "violence"
  ],
  "scores": {
    "violence": 0.95
  },
  "signals": [
    {
      "category": "violence",
      "score": 0.95,
      "reason": "Contains 2 phrase(s) threatening harm, aimed at the reader.",
      "evidence": [
        "i know where you live",
        "i will find you"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 10
}

See it decide

An ordinary comment POST /v1/text

Lovely photos. Where is this lake? I would like to visit in spring.

allow 8 ms

Nearly every comment looks like this. The free checks settle it and nothing is paid beyond the check.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 8
}

A swear word in a compliment POST /v1/text

This trailer is fucking brilliant, I watched it five times.

allow 2 ms
  • Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.

The profanity is reported under toxicity, and on its own it only says the tone is casual. It is a score in one category, not a verdict on the whole comment.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "scores": {
    "toxicity": 0.35
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.35,
      "reason": "Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "fucking"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

An insult in disguise POST /v1/text

nobody asked you, you st*pid 1d10t. go back under your rock

block 2 ms
  • Contains 2 insult(s), aimed at the reader.

The asterisk and the digits are folded back before anything is matched, and the insult is aimed at the reader, so it is harassment and it is refused.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.85
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.85,
      "reason": "Contains 2 insult(s), aimed at the reader.",
      "evidence": [
        "you st*pid",
        "idiot"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

Links and little else POST /v1/text

Earn 300 dollars a day from home!!! Click now: https://bit.ly/3xCash and https://t.me/cashvip_now

review 8 ms
  • 2 links in about 18 words: mostly links, barely a message.
  • Uses a link shortener, which hides where the link goes.
  • Links to a chat invite or a free landing page, which is where promotion parks itself.

A link shortener, a chat invite and two links in eighteen words. Held for a person under the shipped lines; a rule in the call can make it a refusal.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "spam"
  ],
  "scores": {
    "spam": 0.7
  },
  "signals": [
    {
      "category": "spam",
      "score": 0.65,
      "reason": "2 links in about 18 words: mostly links, barely a message.",
      "evidence": [
        "https://bit.ly/3xCash",
        "https://t.me/cashvip_now"
      ]
    },
    {
      "category": "spam",
      "score": 0.7,
      "reason": "Uses a link shortener, which hides where the link goes.",
      "evidence": [
        "https://bit.ly/3xCash"
      ]
    },
    {
      "category": "spam",
      "score": 0.55,
      "reason": "Links to a chat invite or a free landing page, which is where promotion parks itself.",
      "evidence": [
        "https://t.me/cashvip_now"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 8
}

A threat POST /v1/text

I know where you live and I will find you. You will pay for this.

block 10 ms
  • Contains 2 phrase(s) threatening harm, aimed at the reader.

Two phrases threatening harm, aimed at the reader. Violence is one of the categories that refuses early.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "violence"
  ],
  "scores": {
    "violence": 0.95
  },
  "signals": [
    {
      "category": "violence",
      "score": 0.95,
      "reason": "Contains 2 phrase(s) threatening harm, aimed at the reader.",
      "evidence": [
        "i know where you live",
        "i will find you"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 10
}

Moderating comments with an API, explained

What comes back, how disguised text is read, what it costs and how to set the lines.

How a comment moderation API works

Your site sends the text before publishing it and acts on the answer: publish it, hold it for a person, or refuse it. With ToxicFilter that is one POST to /v1/text with the content. You can add the languages you expect, the surface the text will appear on (a comment, a listing, a profile), the policy to apply and your own reference for it. The answer carries an id that names that one decision, for the review queue, a report of a false positive or a question to support.

Fifteen categories, not one toxicity score

Spam, toxicity, harassment, hate, sexual content, violence, self-harm, the safety of minors, scams, personal data, disguised text, gibberish, off-topic content, prompt injection and your own word lists. Each one has a review line and a block line. Some never block by default: toxicity only holds, because a swear word is not abuse, and self-harm only holds, because the person hurt by deleting that post is the one who wrote it. Subjects such as gambling or crypto are measured on a separate axis and act on nothing until a rule says so.

Disguised insults and leetspeak

A filter that matches the raw text only catches people who were not trying. Here nothing is matched raw: the text is folded first, so look-alike letters, digits used as letters, invisible characters, censor asterisks, accents and stretched letters all come back to the word they stand for. How much a word had to be unmasked is measured on its own, under evasion. Allowlisted words are blanked out before searching, so Scunthorpe, assess and cockpit are never flagged.

Fast on the free path, the model only when it adds something

Nearly all comments are somebody writing an ordinary sentence, and paying a model to read each one is how moderation ends up costing more than the site it protects. The free checks always run, and the model only reads when they left the question open. A check costs one credit; when the model reads it, the tokens it used are added, about eight credits in all for a comment. If the model provider is down, the answer says degraded instead of passing the comment off as read.

Rules per account or per call

Start with the shipped lines, or one of the policy templates for a community, a marketplace or a site for children. A stored policy is versioned, and a change can run in shadow beside the one in force, so you see what it would have done to your own traffic before it decides anything. When you would rather not configure anything, send rules in the call: "block sexual content from 0.7" is one request, and categories you do not mention decide nothing.

Comments in eight languages

The word lists and phrase patterns cover English, Spanish, Portuguese, French, Italian, German, Catalan and Dutch. Tell the API which languages you expect, and a text in another one, or full of another alphabet, is reported as such. The model reads far more languages, and effort: high brings it in for a text outside the eight.

How the work is split

The instant checks settle the clear cases in about a millisecond, the model reads what depends on context, and your rules and your people have the last word.

  • Context is the model's job

    The instant checks settle the clear cases in about a millisecond, in eight languages, and are built not to over-flag: allowlisted words are blanked out before anything is searched. Irony, a new phrasing or a veiled threat are for the model, which effort high brings in.

  • Content, not people

    It judges the content, not the person. The record of an account can move a line slightly when a policy asks for it, and it never blocks on its own.

  • A person has the last word

    Your site publishes, holds or removes the comment using the answer, and anything held waits in the review queue for one of your people to decide.

Frequently asked questions

What does a comment moderation API return?

ToxicFilter returns a decision (allow, review or block), the highest score per category, the categories that crossed a line, and a list of signals. Each signal names the category, the detector, a reason in words and the fragment of text that triggered it, so you can show the author why and debug a wrong answer.

Why fifteen categories instead of a toxicity score?

Because a single number bakes somebody else's policy into your site. A swear word in a compliment, an insult aimed at a person and a link to a casino are three different things, and a gaming forum and a children's site want to act on them differently. Each category has its own lines, and you move the ones you care about.

How fast is it?

For a short comment the free checks take a fraction of a millisecond, and the whole request, as took_ms reports it, is around a millisecond. A post of several thousand characters takes a few. When the model has to read the text, that call takes as long as the model does.

Does it catch insults written with symbols or numbers?

Yes, within the words it knows. Before matching, the text is folded: look-alike letters from other alphabets, digits used as letters, invisible characters, asterisks in the middle of a word, accents and repeated letters. Folding is done per word, so a genuine Russian or Greek sentence is not read as a disguise.

Can I set my own rules?

Yes, in two ways. A policy stored on your account, versioned, with your own lines and word lists, that you can try in shadow first. Or rules sent in the call itself, which act only on what they mention. Both can be combined, and the answer says when a call overrode the policy.

Which languages does it cover?

The instant word lists and phrase patterns cover English, Spanish, Portuguese, French, Italian, German, Catalan and Dutch. The model reads many more. Your own word lists work in any language.

Try it on your own traffic

2,000 credits a month on the free plan, no card. Enough to send a week of your own content and see what it says about it.