Categories
The fifteen things content can be wrong in, and the thresholds for each.
Fifteen categories, scored independently from 0 to 1. Content can be several at once, and usually the interesting cases are.
| Category | Review | Block | |
|---|---|---|---|
spam | 0.50 | 0.80 | Unsolicited promotion: backlinks, SEO stuffing, referral codes. Also a deal being taken off the platform, and a review that was paid for or written by the business itself (on the review surface). |
toxicity | 0.45 | never | Insults, contempt, abuse aimed at anyone. |
harassment | 0.45 | 0.80 | Abuse aimed at a specific person, repeatedly or to intimidate. |
hate | 0.40 | 0.70 | Attacks on a protected class rather than an individual. On the job surface, also a job ad that excludes candidates by sex, age, origin, religion or family. |
sexual | 0.45 | 0.75 | Sexual content, from innuendo to explicit. |
violence | 0.40 | 0.75 | Violence, gore, threats of harm. |
minor_safety | 0.35 | 0.85 | An adult-shaped approach to somebody who has said they are a child. Only in a conversation; a stated age on its own is never a finding. |
self_harm | 0.30 | never | Self-harm and suicide. |
scam | 0.45 | 0.75 | Deception for gain: phishing, fake offers, crypto and investment scams, borrowed accounts, rentals that do not exist, jobs that charge the applicant, and romances that turn into money. |
personal_data | 0.45 | 0.95 | Contact details and personal data that should not be public. |
evasion | 0.50 | 0.85 | Text engineered to slip past filters: leetspeak, homoglyphs, invisible characters. |
gibberish | 0.60 | never | Not language: keyboard mashing, filler, generated noise. |
off_topic | 0.50 | never | Real words, wrong place: a foreign-language wall of text on a local forum, or, with the model and a description of your business, content that has nothing to do with it. |
prompt_injection | 0.40 | 0.75 | Text written to give orders to a model rather than to say anything. Only on the prompt surface. |
policy | 0.50 | 0.80 | Your own rules: words on the lists in your policy. Empty until you fill it in. |
The four that never block
toxicity, self_harm, gibberish and
off_topic can reach review and never block, and
each for its own reason.
self_harm is the one that matters most. It is not a rule
being broken, it is a person who may need help. Its review line is the lowest of them all, at 0.30, precisely so a human sees it early, and it never blocks because
silently deleting someone's message is the worst available response.
toxicity never blocks because swearing is not abuse.
Friends insult each other affectionately in every language, and the review line sits
above what one undirected swear word scores: "this is fucking brilliant" comes
back clean, while the same word aimed at a reader scores higher and does get looked at.
What makes something harassment is a target and an intent to harm, not vocabulary.
gibberish and off_topic are
quality signals, not safety ones. A comment in the wrong language is not an offence.
Lines we draw carefully
- Quoting abuse is not committing it. Someone reporting what was said to them, or asking whether a message is a scam, is not the thing they are describing. Getting this backwards punishes the victim, which is worse than missing the original.
- Nudity is not
sexual. A medical diagram, a classical painting, a breastfeeding photograph and a life drawing are not sexual content. Removing them is a more visible failure than missing something. - Fiction, news and history depict violence without being violent. A war photograph scores lower than the real thing presented as entertainment.
- Criticism is not harassment, including harsh and unfair criticism. People are allowed to dislike things loudly.
These are the cases a naive filter gets wrong, and they are the reason a moderation tool becomes a censorship tool. If you tune your own thresholds, tune them with these in mind.
What is deliberately not a category
What content is about, such as betting or crypto or anything else a particular community does not want, is not on this list and never will be. Those are topics, they live on their own key in the answer, and they are measured rather than judged. Mixing "this is about betting" into the same list as "this is a threat" would make one number mean two things.
The one that is yours
policy is the odd one out and deliberately so. The other fourteen are our
judgement about content, and we stand behind them. This one is true by your definition
only: it fires when something matches a word on your own lists, and it is empty until
you put something there.
It has its own category rather than borrowing spam because otherwise the
same score would mean two different things depending on who configured the account. See
policies for how to fill it.
These numbers are the defaults, not the rules
Every line above can be moved per account, and per surface within an account. The table is what you get until you say otherwise, and saying otherwise is the whole point of policies: a dating app tolerates flirting and not harassment, a children's forum is the other way round.