What makes a message spam
Most spam reads like a sentence. "Free followers, tap the link" is grammatical, polite and contains no word worth putting on a list. What gives it away is its shape: a message that is mostly links, an address that hides where it goes, a link placed first and filler after it, and the same text arriving dozens of times. ToxicFilter scores each of those as a separate finding, with a reason in words, so you can see which one decided and act on the ones that matter to your site.
How to detect link spam
Counting links is wrong on its own, because one link is how people answer questions. So each call reads every link, including bare domains and the shorteners and chat invites written with no scheme (bit.ly/abc, t.me/abc, discord.gg/abc), and scores what it finds. Several links in fewer than twenty-five words per link start at 0.65 for two and reach 0.75 for three. A message that opens with a link is 0.75. A shortener is 0.70. A chat invite or a free landing page is 0.55. A domain written with spaces round the dot (growfast . xyz) is filed under evasion at 0.65, because it is a statement about intent whatever the link turns out to be.
How to stop repeated messages and spam campaigns
No classifier reading one message can see a campaign, because the campaign is not in the message: it is the same message arriving forty times. ToxicFilter counts each text against the recent calls of the same account and project over the last fifteen minutes, and only for texts of about forty characters or more, so two hundred people writing "thanks!" stay a forum. A fingerprint of each text catches the copies rewritten with a link swapped or two words changed. These counts come from earlier calls, so a single example on a page cannot show them; they appear in the answer as soon as the second copy arrives.
Floods and gibberish
A message that is one word over and over, a key held down, or a wall of emoji scores on spam. Text that does not read as language at all, with almost no vowels and runs of consonants, scores on gibberish, which holds for review and never refuses. Gibberish is only judged on Latin-script text, because vowel counts mean nothing in Chinese or Arabic.
What a repeat costs
A spam campaign is the same message hundreds of times, and the copies are answered from a cache of verdicts kept per account and project. The count of copies is part of what the cache looks up, in steps, so a copy is judged again only when the count has grown enough to change the answer. A cached call is billed one credit, the price of a check, whatever the first one cost. None of the checks on this page uses the model: they are free detectors that settle a comment in about a millisecond.
Choosing the spam thresholds
The shipped lines review spam from 0.50 and block from 0.80, so on their own a shortened link, a message opening with a link or three short links are held for a person rather than refused. The contact form template moves the lines to 0.35 and 0.65, and those same messages are refused. Start from the one that matches the cost of a false positive on your site, then run a second policy in parallel on your own traffic to see where the two would disagree before you switch.