Moderating Comments in Plain PHP with the ToxicFilter SDK
No framework: a comment wall in plain PHP that publishes, holds or refuses each comment, tells the author why, and publishes held ones when a person approves them.
A comment box is the first place spam and abuse arrive on a site, and the last place anybody wants to read every entry by hand. This post builds a small comment wall in plain PHP, with no framework and moderates it with ToxicFilter: every comment is checked before it is shown, and one of three things happens.
- Allow: it is published at once. That is nearly all of them.
- Review: it is held, and a person decides in your ToxicFilter dashboard. A signed webhook tells the app, which publishes it or drops it.
- Block: it is refused, and the author is told why, in words.
Three outcomes and not two on purpose. Forced to choose between publishing and deleting, a strict threshold deletes real comments and a lenient one publishes the abuse; the middle outcome is where the uncertain cases wait for a person. The whole app is in toxicfilter/php-example.
This is one of four posts building the same app with a different client: Flask, Express and Laravel with Laratox. Each one is its own repository, ready to clone or to use as a template: php-example, flask-example, express-example and laravel-example.
Install and configure
The client is one Composer package with no dependencies beyond ext-curl:
composer require edulazaro/toxicfilter-sdk
Create a key in your dashboard; the free plan is enough. A tf_test_ key costs nothing and runs every free check, which is what settles most comments, but it never asks the model: use a live key to see what the model adds. The key is a credential for the whole account, so it lives in the environment, on the server, and never in a page.
TOXICFILTER_KEY=tf_test_... TOXICFILTER_WEBHOOK_SECRET=whsec_... php -S localhost:8000 -t public
The comments are kept in a JSON file (src/Comments.php), so there is no database to set up. Swap it for your own storage: nothing in the moderation depends on it.
Moderating a comment
When the form is posted, the comment is checked before it is stored:
$id = $comments->nextId();
$tf = new Client(getenv('TOXICFILTER_KEY'));
try {
$verdict = $tf->text($body, [
'surface' => 'comment',
'reference' => "comment_{$id}", // how the webhook finds this comment later
]);
} catch (ApiError $e) {
// Nobody could judge it: hold it rather than publish it unread.
$comments->add($id, $name, $body, 'held', null);
header('Location: /?held=1', true, 303);
exit;
}
if ($verdict->blocked()) {
$error = $verdict->reason() ?? 'This comment cannot be published.';
} else {
$comments->add($id, $name, $body, $verdict->needsReview() ? 'held' : 'published', $verdict->id());
header('Location: /' . ($verdict->needsReview() ? '?held=1' : ''), true, 303);
exit;
}
}
}
Four things in those lines are worth noticing.
surfacesays where the text appears. Policies can treat a comment, a listing and a profile differently, and this is how ToxicFilter knows which it is.referenceis the app's own id for the comment. It comes back with the verdict and with every webhook about it, so the app never has to keep ToxicFilter's ids to find its own comments.- There is no "is it toxic" boolean.
blocked(),needsReview()andallowed()are the three answers, and the verdict also carries the categories that crossed a line, every score and the evidence, for when you want more than the decision. - If the API cannot answer at all, the comment is held, not published. Which way to fail is a decision for each site; for a comment box, holding is the safe one.
Telling the author why
A refused comment with no reason is what makes moderation feel arbitrary. Every verdict carries its reasons in words written to be shown to the person who wrote it: reason() is the first, reasons() all of them. The form shows it under the comment, escaped like everything else the page prints:
<?php if ($error): ?><p class="error"><?= $e($error) ?></p><?php endif ?>
When a person decides: the webhook
A held comment waits in your review queue. When somebody approves or rejects it there, ToxicFilter sends moderation.resolved to your endpoint, signed. The handler checks the signature, finds the comment by its reference, and publishes it or drops it:
if ($method === 'POST' && $path === '/webhooks/toxicfilter') {
$event = Webhooks::event(
file_get_contents('php://input'), // the RAW body, before any parsing
$_SERVER['HTTP_X_TOXICFILTER_SIGNATURE'] ?? '',
getenv('TOXICFILTER_WEBHOOK_SECRET') ?: '',
);
if ($event === null) {
http_response_code(400);
exit;
}
if ($event['event'] === 'moderation.resolved' && preg_match('/^comment_(\d+)$/', $event['data']['reference'] ?? '', $m)) {
$event['data']['action'] === 'approved'
? $comments->publish((int) $m[1])
: $comments->remove((int) $m[1]);
}
http_response_code(204);
exit;
}
Webhooks::event() returns null for anything that does not verify, including a missing header, so an unsigned request is a 400 and never reaches the comments.
To try it on your machine, expose the app with a tunnel (cloudflared tunnel --url http://localhost:8000, for instance), add the tunnel's address followed by /webhooks/toxicfilter as an endpoint in Webhooks, and put the signing secret it shows you in TOXICFILTER_WEBHOOK_SECRET. Then post a comment that lands in review, approve it from the review queue, and watch it appear.
Two details make the handler correct rather than merely working. The signature is checked over the raw body, because a body parsed and encoded again is a different string and would never verify. And the comment is found by its reference, the id the app sent with the check, because the webhook carries the verdict and never the comment: ToxicFilter does not keep what it moderates.
What the client does for you
The client retries a rate limit (429) and a server error (5xx) with a growing wait, and never retries a QuotaExhausted (402): one means "try again in a moment", the other "this account is out of credits", and the right behaviour is opposite. Every call carries an idempotency key, so a retry after a timeout is judged and billed once. And it never reads anything but a real answer as allow: a proxy's error page or an empty body is an error, not a verdict. The failures are classes you can catch one by one; all of them extend ApiError, which is what the wall catches.
The whole client is described in the PHP SDK docs.
Where to go from here
This wall uses ToxicFilter at its simplest: no policy of its own, one site, and only text. Each of those can grow when you need it to.
- Your own rules: a policy moves the thresholds per category, adds your own banned words, or measures subjects like gambling or crypto that are not harmful but may not belong on your site. Name it in the call.
- Several sites in one account: give each one a project, with its own activity, review queue and webhooks.
- More than text: the same client checks images, usernames, whole signups and conversations, where a pile-on or an approach that no single message shows becomes visible.
Keep reading
Moderating Comments in Express with the JavaScript SDK
Build a comment wall in Express that publishes, holds or refuses each comment with its reason, keeps the key o...
Moderating Comments in Flask with the Python SDK
Build a comment wall in Flask that publishes, holds or refuses each comment with its reason, verifies the revi...
Moderating Comments in Laravel with Laratox
A validation rule, a facade and a fake: build a comment wall in Laravel that publishes, holds or refuses each...