SalArt-VQA benchmark reveals VLMs fail at understanding image artifacts
Top models detect artifacts 99% of the time but only 53% truly understand them
Researchers introduce SalArt-VQA, a diagnostic benchmark to test whether vision-language models (VLMs) truly understand artifacts in AI-generated images. With 950 images and 3,681 human-authored questions across four types, they tested 20 VLMs. The best model reached 99.37% detection recall but answered all artifact questions correctly on only 53.26% of images, revealing hidden failures in reasoning and calibration.
- Best VLM achieves 99.37% detection recall but only 53.26% correct on all four artifact questions
- Benchmark contains 950 images and 3,681 human-authored questions across four types: presence, localization, grounding, defect identification
- Reveals sensitivity-calibration tradeoff: sensitive models make unsupported claims, conservative models miss real artifacts
Why It Matters
As AI image generation grows, ensuring VLMs truly understand artifacts—not just detect them—is critical for trust and reliability.