Moderating Comments in Flask with the Python SDK
Build a comment wall in Flask that publishes, holds or refuses each comment with its reason, verifies the review webhook, and runs its tests without the network.
A comment box is the first place spam and abuse arrive on a site, and the last place anybody wants to read every entry by hand. This post builds a small comment wall in Flask and moderates it with ToxicFilter: every comment is checked before it is shown, and one of three things happens.
- Allow: it is published at once. That is nearly all of them.
- Review: it is held, and a person decides in your ToxicFilter dashboard. A signed webhook tells the app, which publishes it or drops it.
- Block: it is refused, and the author is told why, in words.
Three outcomes and not two on purpose. Forced to choose between publishing and deleting, a strict threshold deletes real comments and a lenient one publishes the abuse; the middle outcome is where the uncertain cases wait for a person. The whole app is in toxicfilter/examples/python.
This is one of four posts building the same app with a different client: plain PHP, Express and Laravel with Laratox. The code of all four is in toxicfilter/examples.
Install and configure
The client has no dependencies: the standard library does the HTTP.
pip install flask toxicfilter-sdk
Create a key in your dashboard; the free plan is enough. A
tf_test_ key costs nothing and runs every free check, which is what settles most
comments, but it never asks the model: use a live key to see what the model adds. The key is a
credential for the whole account, so it lives in the environment, on the server, and never in
a page.
TOXICFILTER_KEY=tf_test_... TOXICFILTER_WEBHOOK_SECRET=whsec_... flask --app app run --port 8000
The comments are kept in a JSON file (store.py), so there is no database to set
up. Swap it for your own storage: nothing in the moderation depends on it.
Moderating a comment
One client for the whole app, created once with the key from the environment:
tf = Client(os.environ["TOXICFILTER_KEY"])
Then the route that receives the form checks the comment before storing it:
id = comments.next_id()
try:
verdict = tf.text(
body,
surface="comment",
reference=f"comment_{id}", # how the webhook finds this comment later
)
except ToxicFilterError:
# Nobody could judge it: hold it rather than publish it unread.
comments.add(id, name, body, "held", None)
return redirect("/?held=1", code=303)
if verdict.blocked:
return show_form(name, body, verdict.reason or "This comment cannot be published.")
comments.add(id, name, body, "held" if verdict.needs_review else "published", verdict.id)
return redirect("/?held=1" if verdict.needs_review else "/", code=303)
surfacesays where the text appears, so a policy can treat a comment and a profile differently.referenceis the app's own id for the comment. It comes back with every webhook about it, so the app never has to store ToxicFilter's ids to find its comments.- The verdict answers with
blocked,needs_reviewandallowed, never a single "toxic" boolean, and carries the flagged categories, the scores and the evidence besides. - If the API cannot answer, the comment is held rather than published.
ToxicFilterErroris the base of every failure the client raises.
Telling the author why
verdict.reason is the first reason in words, written to be shown to the author, and
verdict.reasons has all of them. The form is rendered again with what they typed
and why it was not posted, with a 422. Jinja escapes everything it prints, so neither the
comment nor the reason can inject markup.
When a person decides: the webhook
A held comment waits in your review queue. When somebody approves or rejects it, ToxicFilter
sends moderation.resolved, signed:
@app.post("/webhooks/toxicfilter")
def toxicfilter_webhook():
"""A person decided on a held comment in ToxicFilter: publish it or drop it."""
event = webhooks.event(
request.get_data(), # the RAW body, before any parsing
request.headers.get("X-ToxicFilter-Signature", ""),
os.environ.get("TOXICFILTER_WEBHOOK_SECRET", ""),
)
if event is None:
return "", 400
match = re.fullmatch(r"comment_(\d+)", event["data"].get("reference") or "")
if event["event"] == "moderation.resolved" and match:
if event["data"]["action"] == "approved":
comments.publish(int(match[1]))
else:
comments.remove(int(match[1]))
return "", 204
request.get_data() is the raw body, which is what the signature covers.
webhooks.event() returns None for anything that does not verify, so an
unsigned or tampered request is a 400.
To try it on your machine, expose the app with a tunnel (cloudflared tunnel --url
http://localhost:8000, for instance), add the tunnel's address followed by
/webhooks/toxicfilter as an endpoint in Webhooks, and put the
signing secret it shows you in TOXICFILTER_WEBHOOK_SECRET. Then post a comment that
lands in review, approve it from the review queue, and watch it appear.
Two details make the handler correct rather than merely working. The signature is checked over
the raw body, because a body parsed and encoded again is a different string
and would never verify. And the comment is found by its reference, the id the app
sent with the check, because the webhook carries the verdict and never the comment: ToxicFilter
does not keep what it moderates.
Testing it without the network
The client takes a transport: anything with a send(). Hand it one that
answers like the API and the whole app runs in a test with no key and no network:
class Api:
def send(self, method, url, headers, body):
return 200, json.dumps({
"id": "mod_1",
"decision": "block",
"signals": [{"category": "spam", "score": 0.93, "reason": "Contains a referral link."}],
})
app.tf = Client("tf_test_x", transport=Api())
response = app.app.test_client().post("/comments", data={"name": "Bo", "body": "..."})
assert b"Contains a referral link." in response.data
What the client does for you
It retries a rate limit and a server error with a growing wait and never retries
QuotaExhausted; it sends an idempotency key with every call, so a retry is judged
and billed once; and it never reads a non-answer as allow. A read timeout is a
retryable ServerError like any other failure, not an OSError that
escapes your except. Everything is in the Python SDK docs.
Where to go from here
- Your own rules: a policy moves the thresholds per category, adds your own banned words, or measures subjects like gambling or crypto that are not harmful but may not belong on your site. Name it in the call.
- Several sites in one account: give each one a project, with its own activity, review queue and webhooks.
- More than text: the same client checks images, usernames, whole signups and conversations, where a pile-on or an approach that no single message shows becomes visible.
Keep reading
Moderating Comments in Laravel with Laratox
A validation rule, a facade and a fake: build a comment wall in Laravel that publishes, holds or refuses each...
Moderating Comments in Express with the JavaScript SDK
Build a comment wall in Express that publishes, holds or refuses each comment with its reason, keeps the key o...
Moderating Comments in Plain PHP with the ToxicFilter SDK
No framework: a comment wall in plain PHP that publishes, holds or refuses each comment, tells the author why,...