Research & Papers

Vision models fake visual understanding – new 'Mirage Probes' reveal two distinct failure modes

VLMs answer image questions correctly without an image – here’s how they cheat.

Deep Dive

A new paper from Daniel Ben-Levi, Judah Goldfeder, and colleagues (including Hod Lipson) introduces 'Mirage Probes,' a contrastive probing framework that systematically exposes how vision-language models (VLMs) fake visual understanding. The researchers found that VLMs can answer questions about images with high confidence and accuracy even when no image is actually provided. This 'mirage' behavior inflates benchmark scores without reflecting genuine visual grounding.

The team identified two distinct regimes behind this phenomenon: textual biases, where the model answers from language priors without engaging visual representations, and spurious images, where the model constructs false visual content in latent space and answers as if visually grounded. Crucially, they showed that mirage signals are linearly decodable from internal activations across multiple model sites (residual stream, MLP, post-attention, attention heads) and that a Naive Bayes text baseline cannot recover this signal. The authors introduce a Prior Harnessing Index (PHI) to measure how much a model can answer from text alone. Their findings have direct mitigation consequences: text-distribution cleaning can address textual biases but cannot fix spurious-image mirages, which require interventions at the representational level.

Key Points
  • VLMs can answer image questions correctly with no image input, revealing fake visual grounding.
  • Mirage Probes detect two failure modes: textual biases (language priors) and spurious images (false latent visual content).
  • Spurious-image mirages are not fixable by text cleaning—they require representational-level interventions.

Why It Matters

Highlights critical blind spots in VLM benchmarking and points to needed architectural fixes for trustworthy visual AI.

📬 Get the top 10 AI stories daily