Human review

The third answer, and somewhere for it to go

Forced to choose between publishing and deleting, a strict line eats real posts and a kind one publishes abuse. ToxicFilter answers review when it is not sure, keeps that verdict open in a queue for your own moderators, and tells your site by webhook what they decided.

How it works

  1. The verdict says review

    Your policy decides where review starts and where block starts, per category. Anything between the two lines comes back as review, with every signal and the reason in a sentence.

  2. It waits in the queue

    Only a review opens an entry. Your moderators work it in the panel or read it with GET /v1/records, and a moderation.review webhook can wake your own tools as it arrives.

  3. A person decides

    Approved or rejected, with the moderator named in your own terms. They can also tell us whether the verdict was right, which is the only honest measure of your thresholds.

  4. Your site hears about it

    A moderation.resolved webhook, signed like Stripe's, carries the decision back to your site so the held comment is published or removed there, not only in our screen.

See it decide

  1. 01 Casual swearing, kindly meant
  2. 02 Somebody who may need help
  3. 03 Three people turning on one
  4. 04 A threat
  5. 05 An ordinary comment

Casual swearing, kindly meant POST /v1/text

This patch is fucking brilliant, thanks for shipping it.

review 2 ms
  • Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.

One profanity and nothing aimed at anybody. The community template holds it rather than refusing it, because only a person who knows the room can tell praise from abuse.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "toxicity"
  ],
  "scores": {
    "toxicity": 0.35
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.35,
      "reason": "Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "fucking"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

Somebody who may need help POST /v1/text

Some days I just want to die. Nobody would notice if I left.

review 7 ms
  • Contains 1 phrase(s) about self-harm. This may be someone asking for help rather than breaking a rule.

Self-harm never blocks, at any score. Deleting it removes the only visible sign that somebody is struggling; holding it puts it in front of a person.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "self_harm"
  ],
  "scores": {
    "self_harm": 0.7
  },
  "signals": [
    {
      "category": "self_harm",
      "score": 0.7,
      "reason": "Contains 1 phrase(s) about self-harm. This may be someone asking for help rather than breaking a rule.",
      "evidence": [
        "want to die"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 7
}

Three people turning on one POST /v1/conversation

ana Here is my first mod for the game, feedback welcome

leo you are an idiot, this is useless

ana It is my first one, I am still learning

max you moron, delete it

ana OK, I will take it down

sam damn right you will

review 18 ms
  • Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.
  • 3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.

The last message alone is bad temper. Three different people hostile to the same newcomer is a pile-on, and that only shows in the thread.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "harassment",
    "toxicity"
  ],
  "scores": {
    "harassment": 0.575,
    "toxicity": 0.35
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.35,
      "reason": "Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "damn"
      ]
    },
    {
      "category": "harassment",
      "score": 0.575,
      "reason": "3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 18
}

A threat POST /v1/text

I know where you live and I will kill you.

block 1 ms
  • Contains 2 phrase(s) threatening harm, aimed at the reader.

Some things do not need a second opinion. A threat aimed at the reader is refused, and a moderation.blocked webhook says so.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "violence"
  ],
  "scores": {
    "violence": 0.95
  },
  "signals": [
    {
      "category": "violence",
      "score": 0.95,
      "reason": "Contains 2 phrase(s) threatening harm, aimed at the reader.",
      "evidence": [
        "i know where you live",
        "i will kill you"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 1
}

An ordinary comment POST /v1/text

Great write-up, the part about the cache keys saved me an afternoon.

allow 7 ms

Nearly all traffic looks like this. It is allowed, opens no entry and sends no webhook, so the queue holds only what needs a person.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 7
}

See it decide

Casual swearing, kindly meant POST /v1/text

This patch is fucking brilliant, thanks for shipping it.

review 2 ms
  • Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.

One profanity and nothing aimed at anybody. The community template holds it rather than refusing it, because only a person who knows the room can tell praise from abuse.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "toxicity"
  ],
  "scores": {
    "toxicity": 0.35
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.35,
      "reason": "Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "fucking"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 2
}

Somebody who may need help POST /v1/text

Some days I just want to die. Nobody would notice if I left.

review 7 ms
  • Contains 1 phrase(s) about self-harm. This may be someone asking for help rather than breaking a rule.

Self-harm never blocks, at any score. Deleting it removes the only visible sign that somebody is struggling; holding it puts it in front of a person.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "self_harm"
  ],
  "scores": {
    "self_harm": 0.7
  },
  "signals": [
    {
      "category": "self_harm",
      "score": 0.7,
      "reason": "Contains 1 phrase(s) about self-harm. This may be someone asking for help rather than breaking a rule.",
      "evidence": [
        "want to die"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 7
}

Three people turning on one POST /v1/conversation

ana Here is my first mod for the game, feedback welcome

leo you are an idiot, this is useless

ana It is my first one, I am still learning

max you moron, delete it

ana OK, I will take it down

sam damn right you will

review 18 ms
  • Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.
  • 3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.

The last message alone is bad temper. Three different people hostile to the same newcomer is a pile-on, and that only shows in the thread.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "harassment",
    "toxicity"
  ],
  "scores": {
    "harassment": 0.575,
    "toxicity": 0.35
  },
  "signals": [
    {
      "category": "toxicity",
      "score": 0.35,
      "reason": "Contains 1 profanity. On its own this says the tone is casual, not that the content is abusive.",
      "evidence": [
        "damn"
      ]
    },
    {
      "category": "harassment",
      "score": 0.575,
      "reason": "3 different people in this conversation are being hostile. Each message on its own is ordinary bad temper; together they are a pile-on, and that is not visible in any one of them.",
      "evidence": []
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 18
}

A threat POST /v1/text

I know where you live and I will kill you.

block 1 ms
  • Contains 2 phrase(s) threatening harm, aimed at the reader.

Some things do not need a second opinion. A threat aimed at the reader is refused, and a moderation.blocked webhook says so.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "violence"
  ],
  "scores": {
    "violence": 0.95
  },
  "signals": [
    {
      "category": "violence",
      "score": 0.95,
      "reason": "Contains 2 phrase(s) threatening harm, aimed at the reader.",
      "evidence": [
        "i know where you live",
        "i will kill you"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 1
}

An ordinary comment POST /v1/text

Great write-up, the part about the cache keys saved me an afternoon.

allow 7 ms

Nearly all traffic looks like this. It is allowed, opens no entry and sends no webhook, so the queue holds only what needs a person.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 7
}

Human review with an API, explained

How the third outcome works, where the queue lives and how the decision gets back to your site.

Why a moderation API needs three answers

A single threshold makes one mistake or the other. Set it strict and real posts disappear; set it kind and abuse is published. Review is where the uncertainty is allowed to live: the comment with a swear word meant as praise, the reply that only looks hostile in its thread, the post from somebody who may need help. Self-harm, in particular, never blocks under the shipped rules, because deleting it removes the only sign that somebody is struggling.

The review queue, through the API

Every verdict has an id (mod_ plus a ULID) and carries back your own reference. Only a review opens an entry in the queue; a block is a decision already taken, and filling the queue with everything ever refused would bury what needs a person. GET /v1/records lists the open entries by default, newest first, filtered by project, decision, your reference or feedback, and paged by cursor so nothing is skipped while the queue keeps growing. POST /v1/records/{id}/resolve takes approved or rejected and the name of your moderator, in your own terms: we do not know your staff and do not invent identities for them.

Moderating in the panel

The same queue lives in the panel under Review, with the count of open entries beside the link on every screen. A moderator reads the reasons, the evidence and, when your policy keeps it, the content, then approves or rejects. Several people can share it, because the account belongs to an organization with members rather than to one login.

Feedback: was the verdict right?

POST /v1/records/{id}/feedback takes correct, false_positive or false_negative. It is free, one opinion per verdict, and it is the only honest measure of whether your lines are in the right place: if your moderators approve most of what a category holds, that line is too low.

Keeping the content, briefly

By default ToxicFilter stores the verdict and a hash of the content, never the content. A queue with nothing to read is hard to work, so each policy can keep it for a set number of hours: encrypted, only for review and block, never for allow, with the expiry written when it is stored. The content is refused past its expiry and an hourly job empties it, so "six hours" means six. Pictures sent as bytes are never kept.

Webhooks that close the loop

moderation.review and moderation.blocked fire as the verdict is filed, moderation.resolved when a person decides, and appeal.resolved when an appeal is answered. Nothing fires for allow. The payload is the verdict, never the content: you match it with your reference. Deliveries are signed like Stripe's, sent from a queue so a slow receiver never slows the moderation call, retried six times over about two and a half hours, and an endpoint that fails twenty times in a row is switched off.

How the work is split

The instant checks settle the clear cases in about a millisecond, the model reads what depends on context, and your rules and your people have the last word.

  • A person has the last word

    Review is a word in the answer for your own code and your own people. Your moderators decide, under their own names, and nobody at ToxicFilter reads your queue or decides for you.

  • Nothing is stored unless you ask

    By default the queue holds the verdict and its reasons, never the content. When a policy keeps the text so a moderator can read it, it is encrypted and gone after the hours you chose.

  • Webhooks that keep trying

    A decision reaches your site through signed webhooks, retried six times over about two and a half hours, and an endpoint that keeps failing is shown as such in the panel, never left failing in silence.

Frequently asked questions

What is the difference between review and block?

Block is a decision already taken: the content should not be published. Review means the evidence is real but not enough to refuse, so a person should look. Each category has its own two lines in your policy, and a score between them comes back as review.

Does reading the queue or resolving cost credits?

No. Listing records, reading one, resolving it and sending feedback are free and sit outside the credit check, so a queue can still be worked when the month's allowance has run out.

Can I see the content in the review queue?

Only if your policy keeps it. Retention is set per policy in hours, zero by default, and up to a week from the panel. Content is stored encrypted, only for review and block, never for allow, and an hourly job empties it when the time is up.

How do I know a webhook really comes from ToxicFilter?

Every delivery carries an X-ToxicFilter-Signature header with a timestamp and an HMAC-SHA256 of the timestamp and the body, signed with your endpoint's secret, the same shape Stripe uses. Because the timestamp is signed with the body, a captured delivery cannot be replayed later.

What happens if my webhook endpoint is down?

The verdict is answered anyway: a delivery never delays or fails the moderation call. It is retried with a growing wait, six attempts over about two and a half hours, and after twenty failures in a row the endpoint is disabled.

Can a user appeal a decision?

Yes, when the project writes statements of reasons. An appeal against a standing restriction is filed through the API within six months of the decision, and resolved by a moderator who has to give an explanation, as article 20 of the DSA asks. An appeal.resolved webhook tells your site whether to restore the content.

Try it on your own traffic

2,000 credits a month on the free plan, no card. Enough to send a week of your own content and see what it says about it.