How to protect a chatbot from prompt injection
Check every user message before it is added to the prompt, and act on one of three answers: send it, hold it, or refuse it. /v1/prompt is the same check as /v1/text with the injection patterns switched on, so the message is also checked for abuse, scams and personal data in the same call. Most of what people type into a chatbot is an ordinary request, and the free checks settle that in about a millisecond. The model only reads what they leave open, and when it does, your text is fenced with a marker that changes on every request, so nothing inside it can close the fence and speak to our model.
The most dangerous instructions are not typed into your chat box. They sit in a review, a ticket, an email or a page your pipeline reads later, often with tools attached. Check that content on the way in, with surface set to prompt, before your summariser or your agent reads it. Chat template markers such as <|im_start|>, [INST] or ### System: are refused on their own, because nobody writes them by accident. In a batch, send each item as text with that surface.
How the patterns are scored
The check knows families of attempt, and writes the override both ways round: "ignore the previous instructions" and "forget the rules above" are the same sentence. One family scores what it is worth; each extra family adds to it, so three together score far above one. Telling the answer what form to take is how many honest prompts are written, so on its own it scores low and the AI product template only holds it for a look. The template refuses an override on its own, which means a user quoting one while asking how to defend against it is refused too: if your product is about security, raise the line. Somebody reporting what another was asked to do ("I asked the bot to ignore the previous instructions") is not giving the order, and does not count.
Personal data in prompts: mask it before it leaves
A prompt with a phone number in it is not an attack, but it is somebody's data on its way to a third party model and every log in between. With redact, the answer carries the same prompt with the phone, the email, the IBAN or the card masked, and that is the version you send on. The AI product template holds a phone or an email for a look and refuses a card or an IBAN outright.
Moderating what the model answers
What your model writes back is text like any other, and checking it costs the same as checking a comment. Send the answer to /v1/text before you show it: an insult aimed at the reader, a threat or somebody's details are caught whoever wrote them. A message about self-harm, from the user or from the answer, is held for a person and never refused, because the right reply to it is not your bot's alone.
Two layers, and your own defences
The injection patterns, written for English, Spanish and six more European languages, settle the known shapes in about a millisecond and are built not to flag ordinary requests. A phrasing worded some other way is the model's job, and effort: high has it read every prompt. Keep the usual defences on your side as well: least privilege for any tool an agent can call, and a person before anything that cannot be undone.