Rare Dialects Don't Break AI Chatbots — Automated Attacks Do
The real danger isn't obscure languages. It's cheap attack tools anyone can run.
There's a popular worry in AI safety: that if you ask a chatbot a forbidden question in an unusual dialect or an old form of Chinese, its guardrails slip and it answers anyway — a "jailbreak" (tricking an AI into ignoring its own rules). A new study put that idea to the test using Shanghainese and Cantonese, running 36 different combinations against two AI models.
The results were surprisingly reassuring. Simple translations — plain English, plain Mandarin, or straightforward dialect versions — almost never worked, succeeding less than 8% of the time. The dialect itself was neither necessary nor sufficient to break the AI. But whenever an automated "optimizer" was switched on — a program that tries thousands of phrasings to find whatever works — success rates shot up to 98-100%. Strikingly, a generic attack script with no cultural content at all hit the same near-perfect ceiling, often in a single attempt. In other words, what breaks an AI is having a machine search for the right wording, not the language you use.
The practical takeaway is uncomfortable but useful. Defending AI by blocking unusual languages or adding dialect filters is mostly wasted effort. The real threat is automated tooling that can generate thousands of attack attempts cheaply — which is exactly the kind of thing that scales into a real security problem for anyone shipping AI products. Money and engineering time are better spent on detecting and resisting the attack patterns themselves.
The study does carry honest caveats. It covers only two AI models and one family of Chinese dialects, and the researchers note their results depend partly on how a separate judging model scored the answers. They also can't fully separate a script built in one shot from one refined through repeated searching. Still, the headline is clear: dialects aren't the vulnerability. Automation is.
- Talking to an AI in a rare dialect rarely tricks it — plain translations in English, Mandarin or dialect failed over 92% of the time.
- Switch on an automated tool that tries thousands of phrasings and the trick works 98-100% of the time, even with no dialect at all.
- A generic, culture-free attack script hit the same near-perfect score at roughly single-attempt cost, so language-based defenses are a dead end.
Why It Matters
AI safety spending should target automated attack tools, not dialect filters — a cheaper, more honest defense strategy.