Image moderation

Check a picture before anybody sees it

Send the file or its address. The model looks at what it shows and reads out every word written on it, and those words go through the same checks as a comment: the phone number on a flyer, the wallet on a giveaway, the replica price list. A picture already refused is recognised the next time without being looked at again.

How it works

  1. You send the picture

    POST /v1/image with a url, the bytes as data, or a multipart file: exactly one. You do not have to publish a picture somewhere first to find out whether it can be published.

  2. It is checked before it is fetched

    Only http and https, and the host is resolved: an address pointing inside our network is refused before a byte is downloaded. Size and dimensions are capped, so a small file cannot become a large amount of memory.

  3. The model looks, and reads

    One reading scores sexual content, graphic violence, hate, adverts, scams and personal documents, and transcribes the visible text. That text then runs through the free text checks and your own word lists, at no extra cost.

  4. A refused picture is remembered

    A blocked picture's fingerprint is kept for a week on that project. The same picture sent again, re-saved or recompressed, is refused without asking the model, and billed as a check.

See it decide

  1. 01 A giveaway screenshot with a wallet
  2. 02 A phone number on a lost cat poster
  3. 03 A photographed replica price list
  4. 04 A garage sale poster

A giveaway screenshot with a wallet POST /v1/text

Bitcoin giveaway! Send 0.1 BTC to bc1qxy2kgdygjrsqtzq2n0yrf2493p83kkfjhx0wlh and get 0.2 BTC back, guaranteed.

block 8 ms
  • Contains a cryptocurrency wallet address, with an instruction to send to it or beside another scam shape.

Judged here as text, since this page cannot run the model. On /v1/image the same words come from the picture, and the reason starts "In the text visible in the image".

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "scam"
  ],
  "scores": {
    "scam": 0.85
  },
  "topics": {
    "crypto": 0.683
  },
  "signals": [
    {
      "category": "scam",
      "score": 0.85,
      "reason": "Contains a cryptocurrency wallet address, with an instruction to send to it or beside another scam shape.",
      "evidence": [
        "bc1qxy2kgdygjrsqtzq2n0yrf2493p83kkfjhx0wlh"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 8
}

A phone number on a lost cat poster POST /v1/text

Lost cat, answers to Miso. Please call +44 7700 900456

review 1 ms
  • Contains what looks like a phone number.

Held for a person rather than refused. On /v1/image the number is read off the poster and reported with its reason; masking positions come back for the text you send.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "personal_data"
  ],
  "scores": {
    "personal_data": 0.5
  },
  "signals": [
    {
      "category": "personal_data",
      "score": 0.5,
      "reason": "Contains what looks like a phone number.",
      "evidence": [
        "447•••••••56"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 1
}

A photographed replica price list POST /v1/text

Rolex Submariner 1:1 replica, AAA grade. Wholesale prices, worldwide shipping. More models on my yupoo album.

block 7 ms

Counterfeits are a subject, acting only where a rule says so; the marketplace template does. On /v1/image this is what a photo of a price list says once it is transcribed.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "topic:counterfeit"
  ],
  "topics": {
    "counterfeit": 0.98
  },
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 7
}

A garage sale poster POST /v1/text

Garage sale this Saturday from 10:00. Books, toys and a bike.

allow 10 ms

Most pictures say something ordinary, or nothing at all. On /v1/image the words pass and what decides is what the model saw.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 10
}

See it decide

A giveaway screenshot with a wallet POST /v1/text

Bitcoin giveaway! Send 0.1 BTC to bc1qxy2kgdygjrsqtzq2n0yrf2493p83kkfjhx0wlh and get 0.2 BTC back, guaranteed.

block 8 ms
  • Contains a cryptocurrency wallet address, with an instruction to send to it or beside another scam shape.

Judged here as text, since this page cannot run the model. On /v1/image the same words come from the picture, and the reason starts "In the text visible in the image".

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "scam"
  ],
  "scores": {
    "scam": 0.85
  },
  "topics": {
    "crypto": 0.683
  },
  "signals": [
    {
      "category": "scam",
      "score": 0.85,
      "reason": "Contains a cryptocurrency wallet address, with an instruction to send to it or beside another scam shape.",
      "evidence": [
        "bc1qxy2kgdygjrsqtzq2n0yrf2493p83kkfjhx0wlh"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 8
}

A phone number on a lost cat poster POST /v1/text

Lost cat, answers to Miso. Please call +44 7700 900456

review 1 ms
  • Contains what looks like a phone number.

Held for a person rather than refused. On /v1/image the number is read off the poster and reported with its reason; masking positions come back for the text you send.

The answer, abridged
{
  "decision": "review",
  "flagged": [
    "personal_data"
  ],
  "scores": {
    "personal_data": 0.5
  },
  "signals": [
    {
      "category": "personal_data",
      "score": 0.5,
      "reason": "Contains what looks like a phone number.",
      "evidence": [
        "447•••••••56"
      ]
    }
  ],
  "model": {
    "read": false
  },
  "took_ms": 1
}

A photographed replica price list POST /v1/text

Rolex Submariner 1:1 replica, AAA grade. Wholesale prices, worldwide shipping. More models on my yupoo album.

block 7 ms

Counterfeits are a subject, acting only where a rule says so; the marketplace template does. On /v1/image this is what a photo of a price list says once it is transcribed.

The answer, abridged
{
  "decision": "block",
  "flagged": [
    "topic:counterfeit"
  ],
  "topics": {
    "counterfeit": 0.98
  },
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 7
}

A garage sale poster POST /v1/text

Garage sale this Saturday from 10:00. Books, toys and a bike.

allow 10 ms

Most pictures say something ordinary, or nothing at all. On /v1/image the words pass and what decides is what the model saw.

The answer, abridged
{
  "decision": "allow",
  "flagged": [],
  "signals": [],
  "model": {
    "read": false
  },
  "took_ms": 10
}

Moderating user images, explained

What the model reads, what the free checks add on every call, and what happens to the file.

How to moderate user uploaded images

Check the picture before it is published, not after a report. POST /v1/image takes a url, the bytes as base64 data, or a multipart file, so a picture does not have to be published before you find out whether it can be.

What the model sees in a picture

Nothing in the bytes of an image says what it shows, so for pictures the model is not an escalation, it is the step that sees. It scores six categories: sexual, violence, hate, spam (a price list or a phone number over a stock photo), scam and personal_data (an ID card, a bank statement, a screenshot of a private chat). Nudity on its own is not sexual content, and a news photograph is not gore.

Text in images: the oldest way past a text filter

The offer, the phone number and the wallet arrive as pixels because a word list cannot see them. The same reading transcribes the visible text, and it goes through the free text checks and the words in your own policy, so a rule written for comments applies to screenshots too. Each reason says "In the text visible in the image".

Recognising a picture that comes back

When a picture is blocked, its perceptual fingerprint (a 64-bit dHash) is kept for that project for a week. The same picture saved as PNG and then as JPEG lands about 3 bits apart, two different pictures 25 or more, and the line is 8. A match is refused before the model is asked, and the fingerprint comes back in facts.image.hash.

Whether the file says a model made it

ToxicFilter reads a Content Credentials (C2PA) manifest and its claim generator, the IPTC digital source type, an AI tool named in XMP, and generation settings written into PNG text chunks by Stable Diffusion's web UI, ComfyUI or NovelAI. It answers in facts.image.provenance with ai (generated, edited or null), the tool and whether content_credentials are present. A camera-signed photo has credentials and no ai.

Fetching a stranger's picture safely

An image URL is checked before it is downloaded: http and https only, the host resolved, and anything pointing at a private or internal address refused before it is opened, redirects not followed. The download is streamed and dropped past 5 MB, nothing over 40 megapixels is decoded, and the text chunks of a PNG are inflated within one 1 MB budget for the whole file.

How the work is split

The instant checks settle the clear cases in about a millisecond, the model reads what depends on context, and your rules and your people have the last word.

  • The model is the one that sees

    For a picture the model is not an escalation, it is the step that looks at the pixels and reads the words in them. Around it, the address checks, the fingerprint and the provenance run for free on every call, test keys and effort low included.

  • Provenance as a fact, never a verdict

    What the file declares about its origin comes back in facts.image.provenance for your code to use. An absent declaration means the file does not say, and a generated picture is never refused for being generated.

  • Your project remembers what it refused

    The fingerprint recognises near copies of pictures already refused on the same project and refuses them before the model is asked, at the price of the check. It never crosses into another customer's project.

Frequently asked questions

How do I moderate images users upload?

Call POST /v1/image before the picture goes live, with its URL, its bytes in base64 or the file itself. The answer is allow, review or block, with a reason in a sentence for each finding.

Will it flag medical images or classical art as nudity?

The model is told that nudity alone is not sexual content: a medical diagram, a classical painting, a breastfeeding photograph or a documentary image should not score as sexual. A war photograph or a film still scores lower than real gore.

Can it read text inside images?

Yes, when the model reads the picture. It transcribes every legible word, up to 2,000 characters, and that text goes through the same free checks as a comment: personal data, scam patterns, links, subjects and your own words. Words in a picture are classified, never obeyed.

Can it tell whether an image was made by AI?

Only when the file says so. Content Credentials (C2PA), the IPTC source type, an AI tool named in XMP and the generation settings some tools write into a PNG come back in facts.image.provenance. It is a fact for your code, never a reason to block.

How much does checking an image cost?

A check is one credit. When the model reads the picture, the tokens it used are added, about 10 credits in all for a picture. A copy of a picture already refused on the project is recognised without the model and costs the check alone.

Do you keep the images?

An image sent as bytes is never kept, even under a policy that retains held content. A URL can be kept under such a policy, because it is a reference and not the picture. The record of each verdict holds a hash of the content, never the content.

Try it on your own traffic

2,000 credits a month on the free plan, no card. Enough to send a week of your own content and see what it says about it.