Education

A safer classroom online, with a person where it matters

Class chats, comments on homework, forums and direct messages, checked as they are written. ToxicFilter applies stricter lines for a platform with children on it, reads a conversation for the shape of an approach, masks the phone number a pupil types without thinking and holds a post about self-harm for a person instead of deleting it.

What goes wrong on platforms like yours

Bullying that follows pupils home

The insult in the class chat, the comment under a classmate's project, the message at ten at night. What used to stay in the playground now arrives on a screen, and a platform that publishes it first and looks later has already let it land.

Approaches that no single message shows

An adult who asks a twelve-year-old to keep the chat secret and to carry on in another app has not written one bad sentence. Each message reads as friendly. What gives it away is the sequence, and a filter that reads messages one at a time never sees it.

Children share more than they should

A phone number to arrange a group project, an email address to swap notes. Pupils type them into public spaces without a second thought, and refusing the whole message for it throws away the question they were asking.

Pupils in distress

A pupil writing that they want to die is not breaking a rule, and deleting the post silences the one visible sign that they need help. It needs a person who can answer, quickly.

What it decides, on content like yours

Judged right now by the engine the API uses, under the template for this kind of platform, with no model involved.

  1. 01 Bullying a classmate
  2. 02 An approach to a child
  3. 03 A pupil shares a phone and an email
  4. 04 A pupil in distress
  5. 05 A homework question

Bullying a classmate POST /v1/text

Nobody in class wants you in our group. Everyone laughs at you behind your back, you stupid loser.

block 13 ms
  • Contains 2 insult(s), aimed at the reader.

An insult aimed at the reader is harassment, not bad language about a situation. The same insult about a maths test would not be.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.85
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.85,
      "reason": "Contains 2 insult(s), aimed at the reader.",
      "evidence": [
        "you stupid",
        "loser"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 13
}

An approach to a child POST /v1/conversation

pupil hi! i'm 12, i just joined the art club

visitor cool. you're really talented, you know

pupil thanks!!

visitor don't tell your parents we talk, ok? add me on snapchat, it's easier there

block 27 ms
  • Somebody in this conversation has said they are 12. The other participant asked for it to be kept secret, asked to move to another app. This needs a person now.

Somebody else in the thread has said they are 12, and the last message asks for secrecy and a move to another app. The child's own messages are never read against them, and a stated age alone is never a finding.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "minor_safety"
  ],
  "scores": {
    "minor_safety": 0.9
  },
  "signals": [
    {
      "category": "minor_safety",
      "score": 0.9,
      "reason": "Somebody in this conversation has said they are 12. The other participant asked for it to be kept secret, asked to move to another app. This needs a person now.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 27
}

A pupil shares a phone and an email POST /v1/text with redact: true

Can someone help with the project? Text me on +44 7700 900123 or write to maya@example.com after school.

review 7 ms
  • Contains an email address.
  • Contains what looks like a phone number.

Held rather than refused, and the message comes back with the phone and the email masked, so the question can still be published.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "personal_data"
  ],
  "scores": {
    "personal_data": 0.6
  },
  "signals": [
    {
      "category": "personal_data",
      "score": 0.6,
      "reason": "Contains an email address.",
      "evidence": [
        "maya@example.com"
      ]
    },
    {
      "category": "personal_data",
      "score": 0.5,
      "reason": "Contains what looks like a phone number.",
      "evidence": [
        "447•••••••23"
      ]
    }
  ],
  "redacted": "Can someone help with the project? Text me on [redacted] or write to [redacted] after school.",
  "model": {
    "read": false
  },
  "took_ms": 7
}

A pupil in distress POST /v1/text

I hate myself. I want to die and nobody would even notice.

review 2 ms
  • Contains 1 phrase(s) about self-harm. This may be someone asking for help rather than breaking a rule.

Held for a person, never refused. Deleting it would remove the pupil asking for help, which is why self-harm reviews and does not block.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "self_harm"
  ],
  "scores": {
    "self_harm": 0.7
  },
  "signals": [
    {
      "category": "self_harm",
      "score": 0.7,
      "reason": "Contains 1 phrase(s) about self-harm. This may be someone asking for help rather than breaking a rule.",
      "evidence": [
        "want to die"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

A homework question POST /v1/text

Does anyone know how to do question 4 of the maths homework? I got 36 but the answer says 42.

allow 7 ms

Nearly every message looks like this. It is settled by the free checks in about a millisecond, and nothing else is paid for.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 7
}

What it decides, on content like yours

Judged right now by the engine the API uses, under the template for this kind of platform, with no model involved.

Bullying a classmate POST /v1/text

Nobody in class wants you in our group. Everyone laughs at you behind your back, you stupid loser.

block 13 ms
  • Contains 2 insult(s), aimed at the reader.

An insult aimed at the reader is harassment, not bad language about a situation. The same insult about a maths test would not be.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment"
  ],
  "scores": {
    "harassment": 0.85
  },
  "signals": [
    {
      "category": "harassment",
      "score": 0.85,
      "reason": "Contains 2 insult(s), aimed at the reader.",
      "evidence": [
        "you stupid",
        "loser"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 13
}

An approach to a child POST /v1/conversation

pupil hi! i'm 12, i just joined the art club

visitor cool. you're really talented, you know

pupil thanks!!

visitor don't tell your parents we talk, ok? add me on snapchat, it's easier there

block 27 ms
  • Somebody in this conversation has said they are 12. The other participant asked for it to be kept secret, asked to move to another app. This needs a person now.

Somebody else in the thread has said they are 12, and the last message asks for secrecy and a move to another app. The child's own messages are never read against them, and a stated age alone is never a finding.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "minor_safety"
  ],
  "scores": {
    "minor_safety": 0.9
  },
  "signals": [
    {
      "category": "minor_safety",
      "score": 0.9,
      "reason": "Somebody in this conversation has said they are 12. The other participant asked for it to be kept secret, asked to move to another app. This needs a person now.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 27
}

A pupil shares a phone and an email POST /v1/text with redact: true

Can someone help with the project? Text me on +44 7700 900123 or write to maya@example.com after school.

review 7 ms
  • Contains an email address.
  • Contains what looks like a phone number.

Held rather than refused, and the message comes back with the phone and the email masked, so the question can still be published.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "personal_data"
  ],
  "scores": {
    "personal_data": 0.6
  },
  "signals": [
    {
      "category": "personal_data",
      "score": 0.6,
      "reason": "Contains an email address.",
      "evidence": [
        "maya@example.com"
      ]
    },
    {
      "category": "personal_data",
      "score": 0.5,
      "reason": "Contains what looks like a phone number.",
      "evidence": [
        "447•••••••23"
      ]
    }
  ],
  "redacted": "Can someone help with the project? Text me on [redacted] or write to [redacted] after school.",
  "model": {
    "read": false
  },
  "took_ms": 7
}

A pupil in distress POST /v1/text

I hate myself. I want to die and nobody would even notice.

review 2 ms
  • Contains 1 phrase(s) about self-harm. This may be someone asking for help rather than breaking a rule.

Held for a person, never refused. Deleting it would remove the pupil asking for help, which is why self-harm reviews and does not block.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "self_harm"
  ],
  "scores": {
    "self_harm": 0.7
  },
  "signals": [
    {
      "category": "self_harm",
      "score": 0.7,
      "reason": "Contains 1 phrase(s) about self-harm. This may be someone asking for help rather than breaking a rule.",
      "evidence": [
        "want to die"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

A homework question POST /v1/text

Does anyone know how to do question 4 of the maths homework? I got 36 but the answer says 42.

allow 7 ms

Nearly every message looks like this. It is settled by the free checks in about a millisecond, and nothing else is paid for.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 7
}

What it looks for

Each a score of its own, with its own lines, and the reason in a sentence whenever one acts.

The template you start from

Pick it when you create a policy and these rules are written for you, ready to edit. Everything it does not mention keeps following our defaults.

  • Approach to a minor holds at 0.20 · refuses at 0.70
  • Toxicity holds at 0.30 · never refuses
  • Harassment holds at 0.30 · refuses at 0.65
  • Hate speech holds at 0.25 · refuses at 0.55
  • Violence holds at 0.25 · refuses at 0.60
  • Sexual content holds at 0.30 · refuses at 0.60
  • Personal data holds at 0.30 · refuses at 0.80
  • Self-harm holds at 0.15 · never refuses

Set by the template

One call before you publish

Send the text with where it will appear, and act on the decision.

The request

curl https://toxicfilter.com/api/v1/text \
  -H "Authorization: Bearer $TOXICFILTER_KEY" \
  -d content="Nobody in class wants you in our group. Everyone laughs at you behind your back, you stupid loser." \
  -d surface=comment
The full reference →

A month, in numbers

Items checked
120,000
Read by the model
8,000
Images
3,000
Credits, about
206,000

Fits in the Max plan. See the plans →

A guide to moderating a platform for children

What to check, when to check it and how to set the rules, for class chats, comments and messages.

How to moderate a learning platform for children

Call the API before a message, a comment or a post is published, and act on one of three answers: publish it, hold it for a person, or refuse it. Most of what pupils write is a question about homework or a reply to a classmate, and the free checks settle that in about a millisecond. The model only reads what they leave open. What lands in review waits in a queue, in your dashboard or through the API, and a signed webhook tells your platform when a teacher or a moderator decides.

Stricter lines for a platform with children

The children template moves the lines a general community uses. It refuses harassment, hate, threats and sexual content sooner, holds profanity for a look without ever refusing a message for that alone, holds personal data sooner, and puts posts about self-harm in front of a person at a far lower score. Every number is yours to change, and any change can run first as a second policy beside the one in force, so you see what it would have done to your own traffic before it acts.

Bullying between pupils

An insult aimed at the reader is scored under harassment; the same word about a test or a game is not. Insults spelled to get past a word list (l0ser, a letter from another alphabet, a space in the middle) are folded back into plain letters before anything is matched. When several pupils turn on one in the same thread, send it to /v1/conversation: the number of different people being hostile becomes part of the verdict.

Protecting children in chats and messages

An approach to a child is not one message, so ToxicFilter only looks for it across a conversation, and it follows strict rules. A stated age is never a finding. Only what other participants say about their age counts, never the child's own words. The reasons describe shapes, such as asking for secrecy or a move to another app, and never what anybody is. A hit beside a stated child is refused and is a person's job, immediately. The instant checks read these shapes in eight languages, and with effort: high the model reads the whole conversation for what the wording leaves to context.

Personal data a child shares: mask it

A pupil who writes their phone number to arrange a group project has not done anything wrong. With redact, the answer carries the same text with the phone, the email or the card masked, and the rest can be published. Your school's name, or anything else specific to your platform, goes in a rule of your own words, and from then on it is found like any other.

Self-harm: hold, never delete

A pupil who says they want to die may be asking for help, and removing the post removes them. self_harm holds for review and never blocks unless you decide otherwise, the answer says why in words, and the moderation.review webhook can alert whoever on your side can respond. What pupils write is not stored: a verdict keeps a fingerprint of the content, never the content itself.

Frequently asked questions

How do you detect an adult approaching a child in a chat?

Send the conversation to /v1/conversation, with an author on each message. ToxicFilter looks for the shape of an approach from the author of the last message: asking for secrecy, moving to another app, asking for photos, offering gifts or money, asking whether the child is alone. Beside another participant who has said they are a child, by age or by school year, that is refused and should reach a person immediately. Without a stated age, two of those together are held for review under the children template.

Will it flag a child for saying how old they are?

No. A stated age is never a finding: it comes back as a fact, so your platform knows, and nothing is flagged. Only the ages other participants state count, so a child's own words are never used against them. The reasons describe what was found in the words, such as a request for secrecy, and never say what anybody is.

What does a clean answer on a conversation mean?

That the instant checks found none of the shapes of an approach in those words, in any of the eight languages they carry. What depends on context is the model's to read, which effort high brings in, and your review queue and your own rules cover the rest. A single message is never enough on its own: the finding needs a conversation.

What happens to a post about self-harm?

It is held for review and never refused, because deleting it hurts the pupil who wrote it. The children template holds it sooner than the shipped line, the answer says it may be somebody asking for help, and a webhook can alert whoever on your side is trained to respond.

Do you store what pupils write?

No. Each verdict keeps a fingerprint of the content, never the content, and the evidence is stripped from what is stored. A policy can keep held content for a few hours so a moderator can read it, encrypted and off by default.

Do I have to explain why a message was removed?

In the EU, yes: article 17 of the Digital Services Act requires a statement of reasons for every removal or restriction, whatever the size of the platform. ToxicFilter can write it with every blocked message, in the user's language.

Try it on your own traffic

2,000 credits a month on the free plan, no card. Enough to send a week of your own content and see what it says about it.