Why a moderation API needs three answers
A single threshold makes one mistake or the other. Set it strict and real posts disappear; set it kind and abuse is published. Review is where the uncertainty is allowed to live: the comment with a swear word meant as praise, the reply that only looks hostile in its thread, the post from somebody who may need help. Self-harm, in particular, never blocks under the shipped rules, because deleting it removes the only sign that somebody is struggling.
The review queue, through the API
Every verdict has an id (mod_ plus a ULID) and carries back your own reference. Only a review opens an entry in the queue; a block is a decision already taken, and filling the queue with everything ever refused would bury what needs a person. GET /v1/records lists the open entries by default, newest first, filtered by project, decision, your reference or feedback, and paged by cursor so nothing is skipped while the queue keeps growing. POST /v1/records/{id}/resolve takes approved or rejected and the name of your moderator, in your own terms: we do not know your staff and do not invent identities for them.
Moderating in the panel
The same queue lives in the panel under Review, with the count of open entries beside the link on every screen. A moderator reads the reasons, the evidence and, when your policy keeps it, the content, then approves or rejects. Several people can share it, because the account belongs to an organization with members rather than to one login.
Feedback: was the verdict right?
POST /v1/records/{id}/feedback takes correct, false_positive or false_negative. It is free, one opinion per verdict, and it is the only honest measure of whether your lines are in the right place: if your moderators approve most of what a category holds, that line is too low.
Keeping the content, briefly
By default ToxicFilter stores the verdict and a hash of the content, never the content. A queue with nothing to read is hard to work, so each policy can keep it for a set number of hours: encrypted, only for review and block, never for allow, with the expiry written when it is stored. The content is refused past its expiry and an hourly job empties it, so "six hours" means six. Pictures sent as bytes are never kept.
Webhooks that close the loop
moderation.review and moderation.blocked fire as the verdict is filed, moderation.resolved when a person decides, and appeal.resolved when an appeal is answered. Nothing fires for allow. The payload is the verdict, never the content: you match it with your reference. Deliveries are signed like Stripe's, sent from a queue so a slow receiver never slows the moderation call, retried six times over about two and a half hours, and an endpoint that fails twenty times in a row is switched off.