Study Finds AI Often Knows the Right Answer — Then Says Something Else
Your AI assistant may already know the truth, but it can't say it.
A researcher at a single-author study tested three well-known AI models that can look at pictures and answer questions about them — LLaVA-1.5-7B, Qwen2.5-VL-7B and InternVL3-8B. These are the same family of tools that power image search, photo description, document scanning and visual chat assistants. The test used POPE, a standard quiz that checks whether an AI hallucinates (makes things up) about what it sees.
The surprising result: when the AI got the answer wrong, the correct answer was already encoded inside the model's middle layers in 68 to 91 percent of cases. The knowledge was there. It just never made it to the output. The researcher then tried to nudge those middle layers by hand — a technique called patching, where you swap a piece of the AI's internal wiring to see what changes. Almost nothing changed. Zero meaningful flips across all three models at the layer level, and on one model, zero out of 12,600 attempts.
Why does that matter to you? Because it reframes AI errors. We tend to assume an AI gives a wrong answer when it doesn't know something. This study suggests the bigger problem is a broken connection between what the AI knows and what it says. That is a wiring problem, not an ignorance problem — and it may be fixable in ways that simply feeding the model more data is not.
The researcher also sorted the mistakes into three buckets: it misread the image, it knew the answer but couldn't express it, or it overrode the image with its own assumptions. A separate tool could tell these apart with better than 60 percent accuracy across all three models. That last category — assuming instead of looking — is the one that produces confident-sounding nonsense. The honest catch: the author explicitly says this is not yet a deployable fix, just evidence that the problem is real.
- Three popular image-reading AIs were wrong for the same reason: the right answer was inside them but never reached the output
- The correct answer was already present in 68 to 91 percent of the models' mistakes
- Researchers can now sort AI errors into three types — misread, knew-but-couldn't-say, and assumed-over-what-it-saw
Why It Matters
AI mistakes may be a fixable wiring problem, not ignorance — meaning future tools could be far more trustworthy.