New study finds VLMs fail multi-script languages like Punjabi
Accuracy delta up to 16% and script consistency as low as 24.8%…
Current vision-language models assume one language equals one script, ignoring billions of users of multi-script languages like Punjabi, Serbian, and Hindi-Urdu. Researchers introduce PuMVR (Punjabi Multimodal Visual Reasoning), the first benchmark to quantify orthographic bias with 375 culturally grounded image-reasoning tasks across Punjabi's three active scripts: Gurmukhi, Shahmukhi, and Roman.
Evaluating 10 state-of-the-art VLMs, they reveal a substantial Script Gap: accuracy deltas reach 16% and Script Consistency Rates (SCR) as low as 24.8%. Visual input boosts absolute performance but doesn't close the gap. Reasoning patterns show limited cross-script transferability, and Chain-of-Thought pathways diverge solely based on script. The authors propose SCR as a core metric for script-agnostic evaluation, challenging current multilingual assessment paradigms.
- PuMVR benchmark: 375 visual reasoning tasks across 3 Punjabi scripts (Gurmukhi, Shahmukhi, Roman).
- Script Gap: accuracy deltas up to 16% between scripts on identical tasks.
- Script Consistency Rate (SCR) as low as 24.8%, highlighting severe bias.
Why It Matters
Exposes hidden bias in multilingual VLMs, demanding script-agnostic evaluation for equitable AI across billions of users.