Technical Oct 2, 2026 · 6 min read

Moderating Comments in Flask with the Python SDK

Build a comment wall in Flask that publishes, holds or refuses each comment with its reason, verifies the review webhook, and runs its tests without the network.

Eduardo Lázaro
Eduardo Lázaro
Founder of ToxicFilter
Moderating Comments in Flask with the Python SDK

A comment box is the first place spam and abuse arrive on a site, and the last place anybody wants to read every entry by hand. This post builds a small comment wall in Flask and moderates it with ToxicFilter: every comment is checked before it is shown, and one of three things happens.

  • Allow: it is published at once. That is nearly all of them.
  • Review: it is held, and a person decides in your ToxicFilter dashboard. A signed webhook tells the app, which publishes it or drops it.
  • Block: it is refused, and the author is told why, in words.

Three outcomes and not two on purpose. Forced to choose between publishing and deleting, a strict threshold deletes real comments and a lenient one publishes the abuse; the middle outcome is where the uncertain cases wait for a person. The whole app is in toxicfilter/examples/python.

This is one of four posts building the same app with a different client: plain PHP, Express and Laravel with Laratox. The code of all four is in toxicfilter/examples.

Install and configure

The client has no dependencies: the standard library does the HTTP.

pip install flask toxicfilter-sdk

Create a key in your dashboard; the free plan is enough. A tf_test_ key costs nothing and runs every free check, which is what settles most comments, but it never asks the model: use a live key to see what the model adds. The key is a credential for the whole account, so it lives in the environment, on the server, and never in a page.

TOXICFILTER_KEY=tf_test_... TOXICFILTER_WEBHOOK_SECRET=whsec_... flask --app app run --port 8000

The comments are kept in a JSON file (store.py), so there is no database to set up. Swap it for your own storage: nothing in the moderation depends on it.

Moderating a comment

One client for the whole app, created once with the key from the environment:

tf = Client(os.environ["TOXICFILTER_KEY"])

Then the route that receives the form checks the comment before storing it:

id = comments.next_id()

try:
    verdict = tf.text(
        body,
        surface="comment",
        reference=f"comment_{id}",  # how the webhook finds this comment later
    )
except ToxicFilterError:
    # Nobody could judge it: hold it rather than publish it unread.
    comments.add(id, name, body, "held", None)
    return redirect("/?held=1", code=303)

if verdict.blocked:
    return show_form(name, body, verdict.reason or "This comment cannot be published.")

comments.add(id, name, body, "held" if verdict.needs_review else "published", verdict.id)

return redirect("/?held=1" if verdict.needs_review else "/", code=303)
  • surface says where the text appears, so a policy can treat a comment and a profile differently.
  • reference is the app's own id for the comment. It comes back with every webhook about it, so the app never has to store ToxicFilter's ids to find its comments.
  • The verdict answers with blocked, needs_review and allowed, never a single "toxic" boolean, and carries the flagged categories, the scores and the evidence besides.
  • If the API cannot answer, the comment is held rather than published. ToxicFilterError is the base of every failure the client raises.

Telling the author why

verdict.reason is the first reason in words, written to be shown to the author, and verdict.reasons has all of them. The form is rendered again with what they typed and why it was not posted, with a 422. Jinja escapes everything it prints, so neither the comment nor the reason can inject markup.

When a person decides: the webhook

A held comment waits in your review queue. When somebody approves or rejects it, ToxicFilter sends moderation.resolved, signed:

@app.post("/webhooks/toxicfilter")
def toxicfilter_webhook():
    """A person decided on a held comment in ToxicFilter: publish it or drop it."""
    event = webhooks.event(
        request.get_data(),  # the RAW body, before any parsing
        request.headers.get("X-ToxicFilter-Signature", ""),
        os.environ.get("TOXICFILTER_WEBHOOK_SECRET", ""),
    )

    if event is None:
        return "", 400

    match = re.fullmatch(r"comment_(\d+)", event["data"].get("reference") or "")

    if event["event"] == "moderation.resolved" and match:
        if event["data"]["action"] == "approved":
            comments.publish(int(match[1]))
        else:
            comments.remove(int(match[1]))

    return "", 204

request.get_data() is the raw body, which is what the signature covers. webhooks.event() returns None for anything that does not verify, so an unsigned or tampered request is a 400.

To try it on your machine, expose the app with a tunnel (cloudflared tunnel --url http://localhost:8000, for instance), add the tunnel's address followed by /webhooks/toxicfilter as an endpoint in Webhooks, and put the signing secret it shows you in TOXICFILTER_WEBHOOK_SECRET. Then post a comment that lands in review, approve it from the review queue, and watch it appear.

Two details make the handler correct rather than merely working. The signature is checked over the raw body, because a body parsed and encoded again is a different string and would never verify. And the comment is found by its reference, the id the app sent with the check, because the webhook carries the verdict and never the comment: ToxicFilter does not keep what it moderates.

Testing it without the network

The client takes a transport: anything with a send(). Hand it one that answers like the API and the whole app runs in a test with no key and no network:

class Api:
    def send(self, method, url, headers, body):
        return 200, json.dumps({
            "id": "mod_1",
            "decision": "block",
            "signals": [{"category": "spam", "score": 0.93, "reason": "Contains a referral link."}],
        })

app.tf = Client("tf_test_x", transport=Api())
response = app.app.test_client().post("/comments", data={"name": "Bo", "body": "..."})
assert b"Contains a referral link." in response.data

What the client does for you

It retries a rate limit and a server error with a growing wait and never retries QuotaExhausted; it sends an idempotency key with every call, so a retry is judged and billed once; and it never reads a non-answer as allow. A read timeout is a retryable ServerError like any other failure, not an OSError that escapes your except. Everything is in the Python SDK docs.

Where to go from here

  • Your own rules: a policy moves the thresholds per category, adds your own banned words, or measures subjects like gambling or crypto that are not harmful but may not belong on your site. Name it in the call.
  • Several sites in one account: give each one a project, with its own activity, review queue and webhooks.
  • More than text: the same client checks images, usernames, whole signups and conversations, where a pile-on or an approach that no single message shows becomes visible.

Put it in front of your real traffic

Allow, review or block, and the reason in words. On the free plan: 2,000 credits a month, no card. A check costs 1 credit, about 8 if the model reads it, about 10 for an image.