New study reveals vision models fail Gestalt tests
Vision models ace benchmarks but flunk human perception tests
A new paper from Sudhanva Manjunath Athreya and Sai Phani Kumar Malladi introduces a behavioral battery to test whether vision models organize visual content like humans do through Gestalt principles—such as grouping by color, shape, or object recognition. The team evaluated 45 models across supervised, self-supervised, contrastive vision-language encoders, open-weight VLMs, and closed foundation models.
The study found that agreement with human responses uncovers perceptual gaps invisible to conventional metrics. Several closed models, despite high benchmark scores, showed poor alignment with human Gestalt judgments. The research provides a reusable framework for auditing whether AI models perceive visual data in ways that align with human audiences, offering a critical check for visualization pipelines.
- Tested 45 models across 5 training families on human-like Gestalt grouping tasks
- Several closed models showed poor human alignment despite high benchmark accuracy
- Introduces a reusable behavioral battery for auditing model-human perceptual alignment
Why It Matters
Highlights critical perceptual gaps in AI vision models that could misalign with human audiences in design and visualization workflows