4 AI models (Claude, GPT-4o, Grok, DeepSeek) share same imperial bias against 91% of world scripts
Tokenizers treat writing systems 31.7x less efficiently; only 9.7% of scripts fully supported.
A new study from Hiroki Fukui reveals that large language models carry forward the digital afterlife of empires, systematically under-supporting the world's writing systems. Using the new Digital Script Representation Index (DSRI), researchers analyzed 300 scripts across seven dimensions. Only 29 (9.7%) are fully supported by modern digital infrastructure, and of 158 living scripts, 38% lack complete support. Tokenizer efficiency — a measure of how many tokens a model needs to represent text — varies by a staggering 31.7x across 45 scripts. The paper introduces a serial mediation model linking imperial intervention to speaker population to web corpus to tokenizer efficiency, finding the direct effect of empire is statistically indistinguishable from zero, suggesting historical inequality is fully mediated through training data.
To test whether these inequalities are model-specific or systemic, Fukui ran 12,000 API calls across four independent LLM families: Anthropic's Claude, OpenAI's GPT-4o, xAI's Grok, and DeepSeek. The results are stunningly convergent: error patterns across these models correlate at Spearman rho of 0.85 to 0.98 (all p < 0.002). A total of 172 script-feature items were answered identically wrong by all four models, with over-attribution errors outnumbering under-recognition by 3.9:1. The feature "used for religion" alone accounted for 43.6% of these convergent errors (4.1x enrichment). Even with religion excluded, the cross-architecture convergence persists (mean rho=0.87) and over-attribution remains at 1.77:1, indicating multi-channeled bias. The findings suggest these biases originate from the shared training corpus — a digital map drawn by colonial and imperial histories — not from individual model architecture choices.
- Only 29 of 300 writing systems (9.7%) have full digital support; 38% of living scripts lack complete tokenizer support
- Tokenizer efficiency varies by 31.7x across languages — vastly different costs for representing the same text
- Four LLM families (Claude, GPT-4o, Grok, DeepSeek) show convergent error patterns (rho=0.85-0.98) with 172 identical wrong answers
Why It Matters
Systemic bias in AI tokenization silently excludes billions of people, perpetuating linguistic and cultural inequality from training data.