AI Plays Games Differently in Different Languages
Your AI helper might be great at English but terrible at Spanish—here's why that matters
New research shows that large language models can play text-based games with different skill depending on the language they use. In a multilingual self-play setup, two copies of the same model competed through different language interfaces, keeping the rules and actions fixed. Testing three open-weight models across eight languages and six games, researchers found systematic differences in win-loss margins, invalid actions, and strategic choices. The findings suggest language can affect different stages of decision-making, and skill gaps are a measurable roadblock to truly multilingual AI.
- The same AI model can play games like chess or poker very differently depending on the language it uses, sometimes winning easily in one language and losing badly in another
- Researchers tested eight languages and found that AI’s strategy, decision-making, and even ability to follow rules varied significantly across languages
- Fixing the language of reasoning (not just the input language) sometimes restored the AI’s performance, suggesting the issue runs deeper than translation
Why It Matters
AI tools may work perfectly in English but fail users in other languages, creating unfair advantages and risks for billions of non-native speakers