Vision language models hallucinate illusions in perfectly normal images
AI mistakes a normal duck for an optical illusion — study reveals systematic overcorrection.
Researcher Tomer Ullman (arXiv, Dec 2024) tested vision-language models on 'illusion-illusions'—images that look like illusions but are actually accurate (e.g., a normal duck, truly crooked lines). Many current vision language systems mistakenly flagged these as illusions, revealing basic processing errors. The study suggests such failures are part of broader issues already discussed in the literature.
- Tomer Ullman (arXiv Dec 2024) presented VLMs with 'illusion-illusions' — images that resemble illusions but are actually accurate (e.g., normal duck, truly different-sized circles).
- Current models like GPT-4V systematically misclassify these as illusions, indicating over-application of illusion heuristics.
- Failures are linked to broader issues in visual reasoning: handling negation, object permanence, and compositional logic.
Why It Matters
Shows vision AI may hallucinate perceptual errors, risking false alarms in real-world visual analysis systems.