Send every comment to the same endpoint, whatever language it is in, with locales set to the languages your site speaks. The word lists for English, Spanish, Portuguese, French, Italian, German, Catalan and Dutch are all searched on every call, so you do not have to detect the language first or route a message to a different model. A short comment the free checks settle comes back in about a millisecond; the model reads it as well when you send "effort": "high".
Which languages the instant checks carry
The vocabulary behind toxicity, harassment, hate, sexual content, violence, self-harm and scams is written in eight languages, and so are the phrase patterns for pile-ons, approaches to children, deals taken off the platform and the scam shapes. Every group must have every language, and a test fails the build when one is missing, because a half-added language looks covered when it is not. The checks that do not depend on words (links, wallet addresses, card and IBAN checksums, repeated messages) work in any language.
Spam in another alphabet
The commonest backlink attack is a wall of Cyrillic dropped into the comments of a site where nobody reads it. With locales set, a text in an alphabet those languages do not use is reported as spam when it has at least fifteen letters of that alphabet and more than one per cent of the text. Links are removed before counting, because a web address is not written in any language and counted as Latin it hid a short Russian bio. A quoted Russian word clears neither gate. Without locales the check stays silent: a multilingual site posting Russian is doing its job.
A message in the wrong language
Fluent English spam on a Spanish forum uses the right alphabet and the wrong language. Over sixty characters, the language check compares the language the text reads as against the ones you declared and reports off_topic only when the winner beats yours by a clear margin, so a Spanish post full of English game titles is left alone. It holds for review and never refuses, since a visitor writing in their own language has done nothing wrong.
Disguised words and genuine foreign script
Nothing is matched against raw text. Look-alike letters are folded word by word, and only where a word mixes alphabets, or sits in a Latin text made of foreign letters that all have a Latin twin, so іdіоt with Cyrillic letters is unmasked while a sentence in Russian keeps its letters. The fold never transliterates to ASCII either, which would erase Chinese, Arabic, Hebrew, Japanese and Korean entirely and make every text in those scripts look like gibberish.
Languages beyond the eight
Outside these eight languages the structural checks still run on every call, and the model reads the language itself. The languages page of the docs sets out what each layer reads. On a site in a language that is not listed, send "effort": "high" for anything that matters, and add the words that matter to you to your own policy lists.