llm-security A field note by Vikrant Sharma
Unicode homoglyphs are still breaking LLM guardrails in 2026
Researchers found that visually identical characters from different scripts let prompts bypass safety filters. The models see different tokens, humans see the same word.
A new paper on arXiv shows that you can still trick LLM safety filters by swapping Latin characters for visually identical Cyrillic or Greek ones. The exploit is embarrassingly simple. Replace the ‘o’ in ‘bomb’ with the Cyrillic ‘о’. To a human, it looks identical. To a tokeniser, it is a completely different byte sequence. The guardrails check for unsafe tokens. The model receives different tokens. The attack works. What surprised me is that this is 2026 and the problem has not been fixed at the tokenisation layer. Most security research assumes you patch the input pipeline once and move on. Unicode normalisation has been solved in web forms for 15 years. But LLM providers are still running models that treat ‘а’ and ‘a’ as unrelated symbols, which means every prompt filter has a blind spot you can drive a truck through. The paper calls this linguistic illegibility. I would call it a failure to apply basic text processing. If your security layer cannot handle homoglyphs, it cannot handle the internet. The fix is not complicated. Normalise to NFC or NFKC before tokenisation. Strip or flag characters outside a known safe set. Treat visually confusable sequences as equivalent for the purpose of content policy. These are standard mitigations in every other domain that processes user text. What this tells me is that LLM security is still being built by people who think in embeddings, not in bytes. The models are getting better at reasoning. The input pipelines are still assuming clean ASCII.
Source: The Implications of Linguistic Illegibility for LLM Security