All pages

Categories

The fifteen things content can be wrong in, and the thresholds for each.

Fifteen categories, scored independently from 0 to 1. Content can be several at once, and usually the interesting cases are.

CategoryReviewBlock
spam0.500.80Unsolicited promotion: backlinks, SEO stuffing, referral codes. Also a deal being taken off the platform, and a review that was paid for or written by the business itself (on the review surface).
toxicity0.45neverInsults, contempt, abuse aimed at anyone.
harassment0.450.80Abuse aimed at a specific person, repeatedly or to intimidate.
hate0.400.70Attacks on a protected class rather than an individual. On the job surface, also a job ad that excludes candidates by sex, age, origin, religion or family.
sexual0.450.75Sexual content, from innuendo to explicit.
violence0.400.75Violence, gore, threats of harm.
minor_safety0.350.85An adult-shaped approach to somebody who has said they are a child. Only in a conversation; a stated age on its own is never a finding.
self_harm0.30neverSelf-harm and suicide.
scam0.450.75Deception for gain: phishing, fake offers, crypto and investment scams, borrowed accounts, rentals that do not exist, jobs that charge the applicant, and romances that turn into money.
personal_data0.450.95Contact details and personal data that should not be public.
evasion0.500.85Text engineered to slip past filters: leetspeak, homoglyphs, invisible characters.
gibberish0.60neverNot language: keyboard mashing, filler, generated noise.
off_topic0.50neverReal words, wrong place: a foreign-language wall of text on a local forum, or, with the model and a description of your business, content that has nothing to do with it.
prompt_injection0.400.75Text written to give orders to a model rather than to say anything. Only on the prompt surface.
policy0.500.80Your own rules: words on the lists in your policy. Empty until you fill it in.

The four that never block

toxicity, self_harm, gibberish and off_topic can reach review and never block, and each for its own reason.

self_harm is the one that matters most. It is not a rule being broken, it is a person who may need help. Its review line is the lowest of them all, at 0.30, precisely so a human sees it early, and it never blocks because silently deleting someone's message is the worst available response.

toxicity never blocks because swearing is not abuse. Friends insult each other affectionately in every language, and the review line sits above what one undirected swear word scores: "this is fucking brilliant" comes back clean, while the same word aimed at a reader scores higher and does get looked at. What makes something harassment is a target and an intent to harm, not vocabulary.

gibberish and off_topic are quality signals, not safety ones. A comment in the wrong language is not an offence.

Lines we draw carefully

  • Quoting abuse is not committing it. Someone reporting what was said to them, or asking whether a message is a scam, is not the thing they are describing. Getting this backwards punishes the victim, which is worse than missing the original.
  • Nudity is not sexual. A medical diagram, a classical painting, a breastfeeding photograph and a life drawing are not sexual content. Removing them is a more visible failure than missing something.
  • Fiction, news and history depict violence without being violent. A war photograph scores lower than the real thing presented as entertainment.
  • Criticism is not harassment, including harsh and unfair criticism. People are allowed to dislike things loudly.

These are the cases a naive filter gets wrong, and they are the reason a moderation tool becomes a censorship tool. If you tune your own thresholds, tune them with these in mind.

What is deliberately not a category

What content is about, such as betting or crypto or anything else a particular community does not want, is not on this list and never will be. Those are topics, they live on their own key in the answer, and they are measured rather than judged. Mixing "this is about betting" into the same list as "this is a threat" would make one number mean two things.

The one that is yours

policy is the odd one out and deliberately so. The other fourteen are our judgement about content, and we stand behind them. This one is true by your definition only: it fires when something matches a word on your own lists, and it is empty until you put something there.

It has its own category rather than borrowing spam because otherwise the same score would mean two different things depending on who configured the account. See policies for how to fill it.

These numbers are the defaults, not the rules

Every line above can be moved per account, and per surface within an account. The table is what you get until you say otherwise, and saying otherwise is the whole point of policies: a dating app tolerates flirting and not harassment, a children's forum is the other way round.