Moderation policies

Your lines, not ours, and tried before they count

One toxicity score for everybody is somebody else's policy. ToxicFilter keeps a line per category, lets you move only the ones you care about, versions every change, and lets you run a new policy beside the live one to see what it would have done before it decides anything.

How it works

  1. Start from a template

    Pick what the policy is for: a community, a marketplace, a dating app, a job board, a children's platform and four more. Its lines are copied into the new policy and are yours to edit from then on.

  2. Move only what you care about

    A policy stores the categories you changed and nothing else. The rest keep following our numbers, so an improvement to them still reaches you. Add your own words to refuse, to look at or to never flag.

  3. Try it in the shadow

    Point the live policy at a second one. Every call is judged under both, only the live one decides, and the dashboard shows where they disagreed on your own traffic. The trial never pays for a second model reading.

  4. Every answer says which version

    The policy's slug and version come back in every answer and stay in every record. Saving a change makes a new version, and nothing judged under the old one is served from the cache again.

See it decide

  1. 01 A swear word in a compliment
  2. 02 Three people on one player
  3. 03 A line the template leaves alone
  4. 04 An ordinary comment

A swear word in a compliment POST /v1/text

This patch is fucking brilliant, thanks for the fix.

review 2 ms
  • Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.

The same sentence is an allow under our default lines: one undirected swear word scores 0.35 on toxicity and our review line is 0.45. The community template holds toxicity from 0.30, so here it reaches a person. Same text, same score, a different policy.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "toxicity"
  ],
  "scores": {
    "toxicity": 0.35
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.35,
      "reason": "Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "fucking"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

Three people on one player POST /v1/conversation

ana you played like shit tonight

ben seriously, you are shit at this game

carl you fucking threw the match

block 16 ms
  • Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.
  • 3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.

Three different people being hostile scores 0.70 on harassment. Under our defaults that is a review, because the block line is 0.80. The community template blocks harassment from 0.65, so here it is refused.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment",
    "toxicity"
  ],
  "scores": {
    "harassment": 0.7,
    "toxicity": 0.55
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.55,
      "reason": "Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "fucking"
      ]
    },
    {
      "category": "harassment",
      "score": 0.7,
      "reason": "3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 16
}

A line the template leaves alone POST /v1/text

Buy cheap followers now at http://bit.ly/x1 http://bit.ly/x2 http://bit.ly/x3 best price click here

review 7 ms
  • 3 links in about 21 words: mostly links, barely a message.
  • Uses a link shortener, which hides where the link goes.

The community template says nothing about spam, so spam follows our lines exactly as if there were no policy: 0.75 is above the review line and under the block line.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "spam"
  ],
  "scores": {
    "spam": 0.75
  },
  "signals": [
    {
      "category": "spam",
      "score": 0.75,
      "reason": "3 links in about 21 words: mostly links, barely a message.",
      "evidence": [
        "http://bit.ly/x1",
        "http://bit.ly/x2",
        "http://bit.ly/x3"
      ]
    },
    {
      "category": "spam",
      "score": 0.7,
      "reason": "Uses a link shortener, which hides where the link goes.",
      "evidence": [
        "http://bit.ly/x1",
        "http://bit.ly/x2",
        "http://bit.ly/x3"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 7
}

An ordinary comment POST /v1/text

Great guide, the second step saved me an hour.

allow 1 ms

Most comments look like this under any policy, and the free checks settle them in about a millisecond.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 1
}

See it decide

A swear word in a compliment POST /v1/text

This patch is fucking brilliant, thanks for the fix.

review 2 ms
  • Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.

The same sentence is an allow under our default lines: one undirected swear word scores 0.35 on toxicity and our review line is 0.45. The community template holds toxicity from 0.30, so here it reaches a person. Same text, same score, a different policy.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "toxicity"
  ],
  "scores": {
    "toxicity": 0.35
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.35,
      "reason": "Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "fucking"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

Three people on one player POST /v1/conversation

ana you played like shit tonight

ben seriously, you are shit at this game

carl you fucking threw the match

block 16 ms
  • Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.
  • 3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.

Three different people being hostile scores 0.70 on harassment. Under our defaults that is a review, because the block line is 0.80. The community template blocks harassment from 0.65, so here it is refused.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment",
    "toxicity"
  ],
  "scores": {
    "harassment": 0.7,
    "toxicity": 0.55
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.55,
      "reason": "Contains 1 profanity, aimed at the reader. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "fucking"
      ]
    },
    {
      "category": "harassment",
      "score": 0.7,
      "reason": "3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 16
}

A line the template leaves alone POST /v1/text

Buy cheap followers now at http://bit.ly/x1 http://bit.ly/x2 http://bit.ly/x3 best price click here

review 7 ms
  • 3 links in about 21 words: mostly links, barely a message.
  • Uses a link shortener, which hides where the link goes.

The community template says nothing about spam, so spam follows our lines exactly as if there were no policy: 0.75 is above the review line and under the block line.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "spam"
  ],
  "scores": {
    "spam": 0.75
  },
  "signals": [
    {
      "category": "spam",
      "score": 0.75,
      "reason": "3 links in about 21 words: mostly links, barely a message.",
      "evidence": [
        "http://bit.ly/x1",
        "http://bit.ly/x2",
        "http://bit.ly/x3"
      ]
    },
    {
      "category": "spam",
      "score": 0.7,
      "reason": "Uses a link shortener, which hides where the link goes.",
      "evidence": [
        "http://bit.ly/x1",
        "http://bit.ly/x2",
        "http://bit.ly/x3"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 7
}

An ordinary comment POST /v1/text

Great guide, the second step saved me an hour.

allow 1 ms

Most comments look like this under any policy, and the free checks settle them in about a millisecond.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 1
}

Moderation policies, explained

What a policy stores, how rules in a call differ, and how to try a change before it decides.

Why one toxicity score is not enough

A gaming forum and a children's platform should not refuse the same sentences. A single number with a single cut-off bakes one site's tolerance into every other site. ToxicFilter scores fifteen categories separately, plus subjects (gambling, crypto, counterfeits) and lead types on their own axes, and each one acts only at the line a policy gives it. The two examples above that change decision are the same text scored the same way: only the lines moved.

How a moderation policy works

A policy stores what you changed and nothing more. The categories you never touched keep following the shipped numbers, so when those improve, the change reaches you too. Switching a category off is a line above 1: it keeps scoring and keeps appearing in the answer, it never acts. Your word lists sit beside the lines: words to refuse, words to look at, and words never to flag, which are blanked out of the text before our own list searches it. A description of what your business does helps the model judge what is off topic for you. Lines can also differ by surface, so a profile and a comment can be held to different numbers under one policy.

Templates per kind of business

Creating a policy starts with what it is for: a community or game, a marketplace, a contact form, a dating app, classifieds, a job board, reviews, an AI product or a platform for children. The template's lines are copied in at that moment and are ordinary rules from then on. Nothing is looked up again later, so editing your policy never surprises anybody else's, and a change to the template does not rewrite yours.

Shadow mode: test a policy on real traffic

The question before every threshold change is what it would have done to last week. Set a second policy as the shadow of the live one. Every call is judged under both, the live one decides alone, and the record keeps both decisions, so the dashboard can show the disagreement over a period of your own traffic. The model's findings are reused as they are and only the free checks run again, so a trial does not double the bill. Images are not re-run, and nothing else changes: same answer, same billing, no extra webhook.

Versions, records and rules in the call

Saving a policy makes a new version. The slug and version come back in every answer and are copied into every record, and the version is part of the cache key, so a verdict reached under the old rules is never served again. A name that does not exist is a 422, never a silent fallback. When you would rather not configure anything first, rules in the request carries the lines for that call alone; it is recorded as inline, which is why a decision that has to be defended later belongs in a named policy.

Reputation, kept on a short leash

A policy can let a person's own record in the same project move the lines a little: at most 0.10, only after 20 verdicts, never across accounts, always reported with the adjustment. It never switches a line on or off, and a good record never loosens minor safety, self-harm, violence or hate, because earning trust first is exactly the pattern those categories exist to see through.

How the work is split

The instant checks settle the clear cases in about a millisecond, the model reads what depends on context, and your rules and your people have the last word.

  • Your rules decide

    A policy decides where every finding acts, category by category, subject by subject and lead type by lead type. What you never touch keeps following the shipped lines, so our improvements still reach you.

  • Try it before it decides

    Shadow mode judges your live traffic under a second policy from the moment the trial starts, reusing the model's findings, so you see what a change would do before it does anything.

  • Every decision can be defended

    Each saved version is kept and copied into every record, so a verdict can be explained months later under the rules in force that day. Rules sent in the call are recorded as inline, for when you would rather not configure anything first.

Frequently asked questions

How do I set different moderation thresholds per category?

Each of the fifteen categories has two lines, review and block, and a policy changes only the ones you add to it. In the panel a rule reads as one line, such as block spam at 0.80; through the API you name the policy with `policy` on any call. A line above 1 is how a category is switched off: it is still scored and reported, it just never acts.

Can I test a moderation policy before turning it on?

Yes. Set another policy as the shadow of the live one. Each call is judged under both, the live one decides, and the record keeps the shadow's decision beside it, so the dashboard shows how many calls the two disagreed on. The model's findings are reused, so a trial does not cost a second reading.

Can I send the rules with the request instead of configuring a policy?

Yes, with `rules` on any endpoint: thresholds, words, subjects, lead types, per-surface lines and the business description for that call. On their own they are exclusive: a category you did not mention decides nothing. Sent with a policy, they are laid over it: only the lines you named change, your words are added to the policy's, and the answer says `overridden: true`.

What happens if I name a policy that does not exist?

The call is refused with a 422 `unknown_policy`. It never falls back to the default, because a typo would otherwise moderate a whole site under rules nobody chose.

Does a user's history change the verdict?

Only if the policy turns reputation on. It then moves the lines by at most 0.10 either way, after 20 verdicts on record, within the same project. A good record never loosens minor safety, self-harm, violence or hate, and the answer always reports the adjustment.

Can each site or app have its own policy?

A policy is either shared by every project of the organization or belongs to one project. A project can name its default policy, and a call can still name another. The word lists blank your allowed words out of the text before our list searches it, so a medical forum or Scunthorpe is not refused for its own vocabulary.

Try it on your own traffic

2,000 credits a month on the free plan, no card. Enough to send a week of your own content and see what it says about it.