PerceptionBench reveals MLLMs fail atomic visual perception, none above 60% accuracy
New benchmark isolates perception from reasoning — results are grim for frontier models.
A large team of researchers led by Zichao Lin has released PerceptionBench, a benchmark specifically designed to isolate and evaluate atomic visual perception in Multimodal Large Language Models (MLLMs). Unlike existing benchmarks that conflate perception with reasoning or domain knowledge, PerceptionBench takes a bottom-up approach: by analyzing failure points across 42 existing benchmarks, the team constructed an error taxonomy defining ten distinct atomic perceptual capabilities. They then built 3,000 verified questions, each targeting a single capability, with difficulty rooted purely in perception rather than reasoning.
Results across 16 frontier MLLMs (including models from OpenAI, Google, and Anthropic) are sobering: no model reached 60% accuracy, and perception-related hallucination emerged as the weakest capability on average. Even models with similar overall scores showed sharply divergent capability profiles, highlighting that current evaluations mask critical gaps. PerceptionBench provides a much-needed capability-level standard for diagnosing where MLLMs truly fail visually, offering researchers a precise tool to improve perception before tackling higher-level reasoning.
- PerceptionBench evaluates 10 atomic visual perception capabilities using 3,000 verified questions with short, unambiguous answers.
- Tested 16 frontier MLLMs — none exceeded 60% accuracy, with perception-related hallucination being the weakest area.
- Derived from analyzing failure points across 42 existing benchmarks to isolate perception from reasoning and domain knowledge.
Why It Matters
Professionals relying on MLLMs for visual tasks need to account for severe perception gaps that current benchmarks hide.