New Method Catches AI When Its Sources Contradict What It Knows
A cheap check spots when AI is torn between two answers before it misleads you.
AI chatbots increasingly answer by first looking things up — a technique called RAG (letting the AI search before it speaks). Usually that helps. But sometimes the documents it finds flatly contradict what the AI already absorbed during training. Picture a friend who insists a restaurant is open until 11, then reads the sign that says 9. A group of researchers wanted to know: can we actually see that internal argument happening, and catch it before a wrong answer reaches us?
Their idea is surprisingly simple. Instead of asking the AI once, ask the same question several times and compare the answers. If the AI confidently knows something, repeated runs land in roughly the same place. If it is torn between a document and its own memory, it wobbles — settling on slightly different answers each time. They call this measurement the Trajectory Variance Score, and it needs as few as two runs of the same question, which makes it cheap to compute.
They tested it on two open AI systems, LLaDA and Dream 7B, and four sets of questions, including general trivia and one built specifically to contain facts that clash with what models memorized. A simple statistical classifier using their score identified conflicts with roughly 70% accuracy. Running the same question five times instead of two nudged accuracy up on one model. Notably, fancier and more complex prediction methods barely beat the simple one.
Why does this matter outside the lab? Nearly every company now attaches AI to its own documents — policies, manuals, customer records. When those documents clash with what the model "thinks" it knows, the AI can quietly blend the two and sound equally confident either way. A cheap wobble-check could flag shaky answers for a human reviewer, or warn a user before they trust a reply. It is not a lie detector: at about 70% accuracy, it is a hint, not a verdict, and it is still research, not a feature you can switch on today.
- AI that looks things up can get confused when the documents it finds disagree with what it already learned.
- By asking the same question twice and comparing answers, researchers detected that confusion about 70% of the time.
- The check is cheap enough to run often, pointing toward future 'this answer might be shaky' warnings in apps.
Why It Matters
Could one day flag when an AI's answer is shaky, so you know when to double-check before trusting it.