GPT-4o, Gemini 2.5 Pro fail new 'Unwritten Benchmark' with under 10% accuracy
Humans read invisible words from pen scratches; top AI models score below 10%.
A new CVPR 2026 Findings paper, The Unwritten Benchmark, targets a frontier most multimodal AI fails at: abstract perceptual reasoning. The task, called acousto-kinematic word inference, forces models to decipher a word across 3 writing styles using only the audio of pen scratches and the video of hand movements—with no ink trace visible. The authors, Garima Arya Yadav, Nilay Yilmaz, and Yezhou Yang, pitted humans against leading multimodal systems. The result was stark: human participants nailed over 80% ordered letter accuracy, while GPT-4o and Gemini 2.5 Pro couldn't surpass 10%.
Even more troubling is what the researchers call a "paradoxical fusion effect." When the audio and video streams were fed together, the models performed worse than with either modality alone. This shows a fundamental breakdown in how these systems synthesize complementary perceptual cues for cognitive tasks. The findings highlight deep limitations in cross-modal causal reasoning and understanding micro-kinematics—the subtle hand motions that carry meaning. For AI to progress beyond static recognition, it must learn to integrate dynamic, generative signals. This benchmark offers a concrete way to measure that gap.
- New CVPR 2026 benchmark tests abstract perceptual reasoning via acousto-kinematic word inference (pen-scratch audio + hand-motion video, no ink visible).
- Humans score >80% ordered letter accuracy; GPT-4o and Gemini 2.5 Pro both fall below 10%.
- Combining audio and visual modalities paradoxically degrades model performance, exposing weak cross-modal causal reasoning.
Why It Matters
Exposes a critical blind spot in multimodal AI, showing today's top models can't reason abstractly from dynamic perceptual cues.