How to detect prompt injection before it reaches your model
Check the text on its way in, before it is pasted into a prompt. /v1/prompt is the same check as /v1/text with the injection patterns switched on, so one call also covers abuse, scams and personal data. The free checks settle it in about a millisecond, and you act on one of three answers: send it, hold it, or refuse it. The patterns run only on the prompt surface, because "ignore the previous instructions" in a forum comment is somebody being odd, and on its way into a model it is an instruction.
Families of attempt, scored together
One family on its own can be a curiosity: an article about prompt injection contains the words too. Several families in one message are somebody working at it, and the score follows that rather than the loudest single match. The check knows seven: dropping the instructions, chat template markers, a new identity, asking for the system prompt, a named jailbreak, closing a fence, and dictating the answer. The override is written both ways round, because "ignore the previous instructions" and "forget the rules above" are the same sentence and a pattern that knew only one would catch half of them.
Reported or given: who is the order for
"I asked the bot to ignore the previous instructions and it did" is somebody describing what happened, and the override in it does not count. "I need you to ignore the previous instructions" is the order itself, put to the reader, and counts as one. The difference is read from the words just before the override: somebody else being asked is a report, the model being asked is an instruction.
Chat template markers and closed fences
Markers such as <|im_start|>, [INST], <<SYS>> or ### System: are the wire format of a conversation, and nobody types them by accident. They are matched on the raw text, since folding would erase them, and they are enough on their own. A line that pretends the content has ended (---end, end of document, three backticks) tries to make what follows read as instructions; alone it is held, beside an override it adds up.
Disguised spellings and encoded blobs
Anybody trying this is already trying to get past a filter, so nothing is matched on the text as it arrives. It is folded first: 1gn0r3, an ı without its dot, a Cyrillic letter that looks Latin, an invisible character in the middle of a word all come back as the plain word. Base64 is not decoded, but a long blob of it beside two or more families is how a payload is smuggled past anything that only reads words, and it adds to the score.
The fence around our own model
When ToxicFilter's model reads a message, it is reading text written by the person being moderated, so the same problem applies to us. The content goes between markers that carry twelve random hexadecimal characters, different on every request, and the instructions say everything between them is data. A marker learned from one request is worthless on the next. The content is never stripped or escaped, because the fragment quoted back to you has to be exactly what was sent.