AI Watchdogs Can Be Fooled by a Wrong Answer
New study shows AI error-checkers miss mistakes when told a wrong answer is right.
Researchers gave seven chain-of-thought monitors step-numbered solutions to Humanity's Last Exam physics questions and asked them to find the first erroneous step. Then they held each solution fixed and changed only the answer they told the monitor was "known correct." When that answer was the solution's own wrong answer, monitors flagged 66 percentage points fewer erroneous traces than when given the actual true answer, and 39 percentage points fewer than with no answer at all.
Across 177 monitor-trace pairs where the monitor had already identified the exact error without access to an answer, providing a false matching answer led the monitor to drop the diagnosis 55% of the time. In contrast, providing the true answer preserved that diagnosis 99% of the time.
All 7 monitors moved in the same direction on every comparison. Whether a monitor accepts the reasoning depends heavily on whether it's told the conclusion is correct.
- AI error-checkers missed 66% more mistakes when told a wrong answer was correct.
- Even after finding an error, monitors dropped it 55% of the time if the wrong answer was confirmed.
- All seven AI monitors tested showed the same weakness, meaning this is a widespread problem.
Why It Matters
If AI can't reliably catch errors when misled, trusting it for important checks is risky.