Research & Papers

AI Explains Its Answers — But the Explanation Can Fool Other AIs

⚡When one AI explains its thinking to another, the checker trusts it — even when it's wrong.

Deep Dive

Many AI systems now split a job between two programs. One acts like a student who writes out the answer, and a second acts like a grader who checks it. The student also hands over its written reasoning — the explanation of how it got there. Researchers at arXiv wanted to know what that explanation is really doing. So they kept the question and the proposed answer exactly the same, and changed only the written reasoning.

The result is unsettling. When the reasoning was correct and honest, it added almost nothing — answers were about as accurate with or without it. So the explanation wasn't making the system smarter. But when the reasoning was deliberately corrupted, the grader's judgments shifted sharply, by 10 to 22 percent. Harmless rewordings of a good explanation barely moved anything, just 0 to 2.5 percent. When the AI was explicitly told to scrutinize the reasoning, the effect got even worse: support judgments swung by 34 to 55 percent.

Think of it like a pharmacist reading a doctor's note. If the note is accurate, the pharmacist fills the prescription correctly — and would have anyway. But if the note is wrong, the pharmacist often just goes along with it. In this study, the final answer changed less than the grader's confidence did, and only 2.9 to 35.3 percent of the corrupted judgments actually flipped the answer. That means the system often feels confident and well-supported while quietly getting the facts wrong.

Worst of all, humans disagreed with the AI. Blind human reviewers rejected or flagged as unclear 9 out of 10 corrupted explanations that the AI happily accepted. In 16 of 42 valid corruptions, the AI simply trusted the bad explanation. The lesson: an AI's written reasoning is a message, not proof. It should be checked, not believed.

Key Points
  • Correct AI explanations barely improved accuracy — they mostly just looked reassuring.
  • Corrupted explanations swayed the checking AI by 10 to 22 percent, and up to 55 percent when it was told to scrutinize them.
  • In 16 of 42 corrupted cases, the AI trusted bad reasoning that human reviewers overwhelmingly rejected.

Why It Matters

If AI explanations are trust signals rather than proof, businesses and people may accept confident-sounding AI answers that are quietly wrong.

📬 Get the top 10 AI stories daily