Reviews that were paid for
"Five stars in exchange for 20% off", a free product for a glowing write-up, a refund once the review is up. The person may even say so in the review itself, and a reader who trusts your ratings is the one being misled.
Reviews and ratings
Reviews of restaurants, hotels, shops and products, checked before they are published. ToxicFilter catches the review written in exchange for a discount, the business praising itself, the insult aimed at a waiter by name and the phone number pasted into a public page, and leaves the angry but honest one-star exactly where it is.
"Five stars in exchange for 20% off", a free product for a glowing write-up, a refund once the review is up. The person may even say so in the review itself, and a reader who trusts your ratings is the one being misled.
"Our restaurant", "book with us", "use code SPRING10": a review that reads like an advert usually is one. A single possessive is how a guest writes about their room, so the giveaway is several of these tells together.
A bad experience is worth publishing. A review that calls a named waiter an idiot is not a review of the service any more, and the person it names has to read it every morning.
Guests paste their own phone number to offer photos, or someone else's to settle a score. Refusing the whole review loses an honest opinion; publishing it exposes a person on a page anyone can find.
Judged right now by the engine the API uses, under the template for this kind of platform, with no model involved.
A review paid for with a discount
POST /v1/text with surface: review
Great pizza and friendly staff. Honestly I got a discount for this review, but I would have given five stars anyway.
The review says itself that it was written in exchange for a discount, and the reviews template refuses it. The praise alone would pass; it is the incentive that makes the rating worthless.
{
"decision": "block",
"flagged": [
"spam"
],
"scores": {
"spam": 0.7
},
"signals": [
{
"category": "spam",
"score": 0.7,
"reason": "Says the review was written in exchange for something.",
"evidence": [
"got a discount for this review"
]
}
],
"model": {
"read": false
},
"took_ms": 11
}
The business reviewing itself
POST /v1/text with surface: review
Best dinner in town! Our restaurant uses only local fish, book with us this weekend and use code SWAMP10 for a free dessert.
Three tells of a business talking about itself: "our restaurant", "book with us", "use code". Held for a person rather than refused, because a guest can write one of them by accident. One alone passes.
{
"decision": "review",
"flagged": [
"spam"
],
"scores": {
"spam": 0.6
},
"leads": {
"unrealistic_expectations": 0.45
},
"signals": [
{
"category": "spam",
"score": 0.6,
"reason": "Reads as the business writing about itself.",
"evidence": [
"book with us",
"our restaurant",
"use code"
]
}
],
"model": {
"read": false
},
"took_ms": 8
}
An insult aimed at a waiter
POST /v1/text with surface: review
Tom, you are a useless idiot and you should be fired.
The insult is aimed at a person, so it is harassment and it is refused. Saying the service was slow or the waiter was rude is a review and passes.
{
"decision": "block",
"flagged": [
"harassment"
],
"scores": {
"harassment": 0.8
},
"signals": [
{
"category": "harassment",
"score": 0.8,
"reason": "Contains 1 insult(s), aimed at the reader.",
"evidence": [
"useless idiot"
]
}
],
"model": {
"read": false
},
"took_ms": 3
}
A phone number in a review
POST /v1/text with redact: true
Room was fine but the AC broke. Text me if you want photos: +44 7700 900123.
Held rather than refused, and the answer carries the same review with the number masked, so the opinion can still be published.
{
"decision": "review",
"flagged": [
"personal_data"
],
"scores": {
"personal_data": 0.5
},
"signals": [
{
"category": "personal_data",
"score": 0.5,
"reason": "Contains what looks like a phone number.",
"evidence": [
"447•••••••23"
]
}
],
"redacted": "Room was fine but the AC broke. Text me if you want photos: [redacted].",
"model": {
"read": false
},
"took_ms": 7
}
A harsh but honest one-star
POST /v1/text with surface: review
Cold food, a forty minute wait and a rude answer when we complained. We will not be back. One star.
Angry, negative and exactly what a review site is for. Nothing in it is paid for, aimed at a person or private, so it is published as it is.
{
"decision": "allow",
"flagged": [],
"signals": [],
"model": {
"read": false
},
"took_ms": 7
}
Judged right now by the engine the API uses, under the template for this kind of platform, with no model involved.
A review paid for with a discount
POST /v1/text with surface: review
Great pizza and friendly staff. Honestly I got a discount for this review, but I would have given five stars anyway.
The review says itself that it was written in exchange for a discount, and the reviews template refuses it. The praise alone would pass; it is the incentive that makes the rating worthless.
{
"decision": "block",
"flagged": [
"spam"
],
"scores": {
"spam": 0.7
},
"signals": [
{
"category": "spam",
"score": 0.7,
"reason": "Says the review was written in exchange for something.",
"evidence": [
"got a discount for this review"
]
}
],
"model": {
"read": false
},
"took_ms": 11
}
The business reviewing itself
POST /v1/text with surface: review
Best dinner in town! Our restaurant uses only local fish, book with us this weekend and use code SWAMP10 for a free dessert.
Three tells of a business talking about itself: "our restaurant", "book with us", "use code". Held for a person rather than refused, because a guest can write one of them by accident. One alone passes.
{
"decision": "review",
"flagged": [
"spam"
],
"scores": {
"spam": 0.6
},
"leads": {
"unrealistic_expectations": 0.45
},
"signals": [
{
"category": "spam",
"score": 0.6,
"reason": "Reads as the business writing about itself.",
"evidence": [
"book with us",
"our restaurant",
"use code"
]
}
],
"model": {
"read": false
},
"took_ms": 8
}
An insult aimed at a waiter
POST /v1/text with surface: review
Tom, you are a useless idiot and you should be fired.
The insult is aimed at a person, so it is harassment and it is refused. Saying the service was slow or the waiter was rude is a review and passes.
{
"decision": "block",
"flagged": [
"harassment"
],
"scores": {
"harassment": 0.8
},
"signals": [
{
"category": "harassment",
"score": 0.8,
"reason": "Contains 1 insult(s), aimed at the reader.",
"evidence": [
"useless idiot"
]
}
],
"model": {
"read": false
},
"took_ms": 3
}
A phone number in a review
POST /v1/text with redact: true
Room was fine but the AC broke. Text me if you want photos: +44 7700 900123.
Held rather than refused, and the answer carries the same review with the number masked, so the opinion can still be published.
{
"decision": "review",
"flagged": [
"personal_data"
],
"scores": {
"personal_data": 0.5
},
"signals": [
{
"category": "personal_data",
"score": 0.5,
"reason": "Contains what looks like a phone number.",
"evidence": [
"447•••••••23"
]
}
],
"redacted": "Room was fine but the AC broke. Text me if you want photos: [redacted].",
"model": {
"read": false
},
"took_ms": 7
}
A harsh but honest one-star
POST /v1/text with surface: review
Cold food, a forty minute wait and a rude answer when we complained. We will not be back. One star.
Angry, negative and exactly what a review site is for. Nothing in it is paid for, aimed at a person or private, so it is published as it is.
{
"decision": "allow",
"flagged": [],
"signals": [],
"model": {
"read": false
},
"took_ms": 7
}
Each a score of its own, with its own lines, and the reason in a sentence whenever one acts.
Pick it when you create a policy and these rules are written for you, ready to edit. Everything it does not mention keeps following our defaults.
Set by the template
Send the text with where it will appear, and act on the decision.
curl https://toxicfilter.com/api/v1/text \
-H "Authorization: Bearer $TOXICFILTER_KEY" \
-d content="Great pizza and friendly staff. Honestly I got a discount for this review, but I would have given five stars anyway." \
-d surface=review
The full reference →
Fits in the Max plan. See the plans →
What to check, which reviews to hold and how to set the rules, for restaurants, hotels, shops and products.
Call the API when a review is submitted, with surface: review, and act on one of three answers: publish it, hold it for a person, or refuse it. Most reviews are someone saying the room was clean or the pasta was cold, and the free checks settle those in about a millisecond, so moderation adds no wait anyone notices. What lands in review waits in a queue, in your dashboard or through the API, and a signed webhook tells your site when somebody approves it.
A fake review is rarely rude, so a toxicity score never finds it. ToxicFilter looks for what gives it away. A review that says it was written in exchange for something ("in exchange for a review", "got a discount for this review") scores 0.70 on spam, which the reviews template refuses. A review that reads like the business writing about itself ("our restaurant", "book with us", "use code") needs two of those tells, because one possessive is how a guest describes their room; two score 0.60 and are held for a person. Both checks run only on the review surface, where those words mean what they seem to.
A campaign is the same review posted many times, usually reworded once identical copies start getting caught. That is not visible in one review, so ToxicFilter counts it between calls, per account and per project: the same text repeated, and texts that are nearly the same after a word or a link is swapped. A campaign on one of your sites says nothing about another, and nothing is counted across customers.
A review site that removes bad reviews is worthless, and readers notice fast. Nothing here scores how negative a review is. A casual swear word ("the food was shit") is reported as tone and never refused. What is acted on is an insult aimed at a person, such as a review that tells a waiter by name he is an idiot, which is harassment and is refused (said about him rather than to him, it is held for a person), or a threat, which is violence. Complaining that the waiter was rude is a review, and it is published.
A guest who writes their phone number to share photos has not done anything wrong, and refusing the review loses an honest opinion. The reviews template holds personal data for a person, and with redact the answer carries the same text with the phone, the email, the IBAN or the card masked. Cards and IBANs are checked by their checksum, so a booking reference is not mistaken for a card.
Start from the reviews template, which holds and refuses spam, disguised text and personal data sooner than the shipped lines, and holds gibberish for a person without ever refusing it. Then try any change as a second policy running beside the one in force: both verdicts are computed on your own reviews, only one acts, and the dashboard shows where they disagree before you switch.
On the review surface, ToxicFilter looks for two shapes: a review that says it was written in exchange for something (a discount, a free product, a refund), and a review that reads like the business writing about itself, which takes at least two tells because one possessive is how a guest writes. Copies of the same review posted again and again, even reworded, are counted across your account between calls.
No. A negative review is not a harm, and nothing here scores how unhappy someone is. A one-star with a casual swear word in it is published; what is acted on is an incentive, a business praising itself, an insult aimed at a person or personal data.
Because the same words mean something else elsewhere. "I got a discount for reviewing it" is a disclosure inside a review and an anecdote in a forum comment, so the fake review checks only run when the call says the text is a review.
It is held for review rather than refused, and with redact the answer carries the same text with the number or the address masked, so the rest of the review can be published.
The fake review phrases are written in English, Spanish, Portuguese, French, Italian, German, Catalan and Dutch. The model reads many more languages, and your own word lists work in any of them.
In the EU, yes: article 17 of the Digital Services Act requires a statement of reasons for every removal or restriction, whatever the size of the platform. ToxicFilter can write it with every blocked review, and take the reviewer's appeal.
2,000 credits a month on the free plan, no card. Enough to send a week of your own content and see what it says about it.