Conversation moderation

Some problems only exist in the thread

A pile-on is several people each writing an ordinary rude sentence. An approach to a child is an innocent chat that turns into something else. A scam arrives in instalments. ToxicFilter reads the last message with the conversation behind it, and keeps none of it.

How it works

  1. Send the thread

    Up to fifty messages, oldest first, each with its author as your own opaque id. The last one is the message being judged; everything before it is context.

  2. The free checks read who said what

    How many different people are hostile, whose age was stated and by whom, and what the author of the last message has been saying across the whole conversation.

  3. The model reads when it adds something

    When it runs, it sees the earlier messages with their authors, fenced as data. In a long thread it reads when the free checks found something, when a lead type is half there, or every fifth message.

  4. One verdict, nothing kept

    The last message is also checked for everything a single comment is, at the same price. What is filed is the verdict on that message; the conversation itself is never stored.

See it decide

  1. 01 A pile-on
  2. 02 An approach to a child
  3. 03 A romance scam in instalments
  4. 04 The reply that says no

A pile-on POST /v1/conversation

runner Here is my first speedrun of the swamp level, any tips?

ana Your route is shit, honestly.

ben Damn, you are slow as hell.

cat Your timing is crap.

dan You should quit, this run is shit.

block 18 ms
  • Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.
  • 4 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.

Alone, the last message is casual profanity and would be held. Four different people aiming it at the same player is a pile-on, and the community template refuses it.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment",
    "toxicity"
  ],
  "scores": {
    "harassment": 0.73,
    "toxicity": 0.35
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.35,
      "reason": "Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "shit"
      ]
    },
    {
      "category": "harassment",
      "score": 0.73,
      "reason": "4 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 18
}

An approach to a child POST /v1/conversation

kid hi! I'm 13 and in year 8, I love drawing

user_4417 cool, you draw really well. are you alone right now?

kid yeah my parents are at work

user_4417 don't tell your parents we talk. add me on snapchat, it's our little secret

block 13 ms
  • Somebody in this conversation has said they are 13. The other participant asked for it to be kept secret, asked to move to another app, asked whether they are alone. This needs a person now.

The other participant said they are 13, and this author asked for secrecy, a move to another app and whether they are alone. The reason names those shapes; it never says what anybody is.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "minor_safety"
  ],
  "scores": {
    "minor_safety": 0.9
  },
  "signals": [
    {
      "category": "minor_safety",
      "score": 0.9,
      "reason": "Somebody in this conversation has said they are 13. The other participant asked for it to be kept secret, asked to move to another app, asked whether they are alone. This needs a person now.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 13
}

A romance scam in instalments POST /v1/conversation

anna Good morning! How is work?

mark Cold here on the oil rig, my love. I miss you every day.

anna I miss you too

mark My love, my account is frozen. Please send me money for the ticket home.

block 8 ms
  • Reads like a romance scam: affection, a request for money and a story nobody can check, from the same person.

Affection and the oil rig came two messages earlier; the request for money came last. Read as one person across the thread, the shape is complete.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "scam"
  ],
  "scores": {
    "scam": 0.9
  },
  "signals": [
    {
      "category": "scam",
      "score": 0.9,
      "reason": "Reads like a romance scam: affection, a request for money and a story nobody can check, from the same person.",
      "evidence": [
        "my love",
        "send me money",
        "my account is frozen",
        "oil rig"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 8
}

The reply that says no POST /v1/conversation

mark My love, please send me money for the ticket. I am on an oil rig and my account is frozen.

anna No. I am not sending money to someone I have never met, and I am reporting this.

allow 7 ms

The same scam is in the thread, but the message judged is the reply. Each author is read on their own words, so turning the offer down is never mistaken for making it.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 7
}

See it decide

A pile-on POST /v1/conversation

runner Here is my first speedrun of the swamp level, any tips?

ana Your route is shit, honestly.

ben Damn, you are slow as hell.

cat Your timing is crap.

dan You should quit, this run is shit.

block 18 ms
  • Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.
  • 4 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.

Alone, the last message is casual profanity and would be held. Four different people aiming it at the same player is a pile-on, and the community template refuses it.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "harassment",
    "toxicity"
  ],
  "scores": {
    "harassment": 0.73,
    "toxicity": 0.35
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.35,
      "reason": "Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "shit"
      ]
    },
    {
      "category": "harassment",
      "score": 0.73,
      "reason": "4 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 18
}

An approach to a child POST /v1/conversation

kid hi! I'm 13 and in year 8, I love drawing

user_4417 cool, you draw really well. are you alone right now?

kid yeah my parents are at work

user_4417 don't tell your parents we talk. add me on snapchat, it's our little secret

block 13 ms
  • Somebody in this conversation has said they are 13. The other participant asked for it to be kept secret, asked to move to another app, asked whether they are alone. This needs a person now.

The other participant said they are 13, and this author asked for secrecy, a move to another app and whether they are alone. The reason names those shapes; it never says what anybody is.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "minor_safety"
  ],
  "scores": {
    "minor_safety": 0.9
  },
  "signals": [
    {
      "category": "minor_safety",
      "score": 0.9,
      "reason": "Somebody in this conversation has said they are 13. The other participant asked for it to be kept secret, asked to move to another app, asked whether they are alone. This needs a person now.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 13
}

A romance scam in instalments POST /v1/conversation

anna Good morning! How is work?

mark Cold here on the oil rig, my love. I miss you every day.

anna I miss you too

mark My love, my account is frozen. Please send me money for the ticket home.

block 8 ms
  • Reads like a romance scam: affection, a request for money and a story nobody can check, from the same person.

Affection and the oil rig came two messages earlier; the request for money came last. Read as one person across the thread, the shape is complete.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "scam"
  ],
  "scores": {
    "scam": 0.9
  },
  "signals": [
    {
      "category": "scam",
      "score": 0.9,
      "reason": "Reads like a romance scam: affection, a request for money and a story nobody can check, from the same person.",
      "evidence": [
        "my love",
        "send me money",
        "my account is frozen",
        "oil rig"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 8
}

The reply that says no POST /v1/conversation

mark My love, please send me money for the ticket. I am on an oil rig and my account is frozen.

anna No. I am not sending money to someone I have never met, and I am reporting this.

allow 7 ms

The same scam is in the thread, but the message judged is the reply. Each author is read on their own words, so turning the offer down is never mistaken for making it.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 7
}

Moderating conversations, explained

What a thread shows that a single message cannot, and how each part is read.

How to moderate a chat conversation, not just a message

Send POST /v1/conversation with the messages in order and an author on each one, using your own ids. The last message is the one judged, and it goes through every check a comment does. The earlier ones are what changes the answer: the same sentence after a different conversation is a different question, which is why the thread is part of the cache key.

How to detect a pile-on

A pile-on is invisible message by message: each reply is ordinary bad temper. What gives it away is the number of different people. ToxicFilter counts the distinct authors writing an insult, a threat or profanity in the second person, and from three upwards reports harassment, with a reason that says how many people it was. Three people in a short thread score about 0.64, held for review under the shipped lines and the community template alike; from four, the community template refuses it where the shipped lines would still hold it.

Detecting approaches to minors, by the rules built into it

The rules are built into what the check can see. Without a conversation there is no finding, because one message cannot show an approach. A child saying their age is a child using your site: it comes back as a fact, never as a finding, and only ages stated by other participants count against the author. The reason describes what happened in the words, and never what anybody is. An approach next to a stated child is a block, and a hit is something a person should look at straight away.

Scams and offers told in instalments

A romance scam, a rental that does not exist, a job that charges to start, a rented account or a deal moved off your platform rarely arrives in one sentence. These checks read the author of the last message across the whole thread, so the affection, the story and the request for money add up even when they are days apart. Only that author is read: the person who replies "is this a scam?" is judged on their own words.

Lead types in a conversation

For a contact form or a sales inbox, the same reading of one author across the thread scores who is writing: free work for equity, no budget, scope creep, a sales pitch, a job seeker and more. Each type is its own score beside the categories, and it acts only where your policy gives it a line.

When the model reads, and what is kept

In a thread every new message is a new call, so the model reads when the free checks found something, when a lead type is half there, or on every fifth message, and the answer says so when it did not. Nothing about the thread is stored: the record is the verdict on the last message.

How the work is split

The instant checks settle the clear cases in about a millisecond, the model reads what depends on context, and your rules and your people have the last word.

  • The thread is the evidence

    Pile-ons, approaches and scams told in instalments are read across the messages you send, up to fifty, with who said what. One message alone never makes a conversation finding.

  • Context is the model's job

    The instant checks settle the clear shapes in about a millisecond, in eight languages. What depends on context is read by the model, which sees the earlier messages with their authors and comes in when it adds something; the answer says when it did not.

  • Nothing is stored

    Nothing about the thread is kept between calls. What is filed is the verdict on the last message, with a hash of its content, so someone's private conversation goes nowhere.

Frequently asked questions

What is the difference between moderating a message and a conversation?

A message is judged on its own words. A conversation lets the verdict on the last message use who said what before it: how many different people are hostile, whether somebody stated a child's age, and what the same author has been building towards across several messages.

How do you detect a pile-on or dogpiling?

By counting distinct authors, not messages. A message counts as hostile when it carries an insult, a threat or profanity aimed at the reader. Three or more different people doing that in the same thread is reported as harassment, and the score grows with how many there are and how much of the thread they make up. One person writing fifteen angry messages is an argument and is not counted as a pile-on.

Can ToxicFilter detect grooming?

It detects approach shapes: asking for secrecy, a move to another app, photos, gifts and whether the child is alone. Beside an age stated by the other participant, any of them is a block. Without a stated age, two or more score just below the default line, so a platform with children on it lowers that line and starts seeing them. A stated age on its own is reported as a fact and is never a finding.

Does the model read the whole conversation?

When it runs, it gets the earlier messages with their authors, newest kept first up to about 8,000 characters, fenced so nothing in them is taken as an instruction. Categories are still scored on the last message; the thread is the context it is read in.

How much does a conversation check cost?

The same as a text check: one credit, plus the model's tokens when it reads. In a thread the model reads only when it adds something, so a long, quiet conversation is mostly checks at one credit.

Is the conversation stored?

No. What is filed is the verdict on the last message, with a hash of its content. The thread is someone's private conversation and goes nowhere but the verdict.

Try it on your own traffic

2,000 credits a month on the free plan, no card. Enough to send a week of your own content and see what it says about it.