Research & Papers

Vision language models hallucinate illusions in perfectly normal images

AI mistakes a normal duck for an optical illusion — study reveals systematic overcorrection.

Deep Dive

Researcher Tomer Ullman (arXiv, Dec 2024) tested vision-language models on 'illusion-illusions'—images that look like illusions but are actually accurate (e.g., a normal duck, truly crooked lines). Many current vision language systems mistakenly flagged these as illusions, revealing basic processing errors. The study suggests such failures are part of broader issues already discussed in the literature.

Key Points
  • Tomer Ullman (arXiv Dec 2024) presented VLMs with 'illusion-illusions' — images that resemble illusions but are actually accurate (e.g., normal duck, truly different-sized circles).
  • Current models like GPT-4V systematically misclassify these as illusions, indicating over-application of illusion heuristics.
  • Failures are linked to broader issues in visual reasoning: handling negation, object permanence, and compositional logic.

Why It Matters

Shows vision AI may hallucinate perceptual errors, risking false alarms in real-world visual analysis systems.

📬 Get the top 10 AI stories daily