Research & Papers

Google's Gemini Spots AI Confusion Cheaper Than GPT-5.2

Multi-step AI systems can misread simple words like 'previous' — and it costs you accuracy.

Deep Dive

Many AI products you already use don't rely on one model. They use a relay: one AI writes a draft, a second AI checks it, and a third AI fixes it up. It's like a writer, an editor and a proofreader working in sequence. This study asked a simple question: when the words 'previous,' 'it' or 'that one' pass down that chain, do all three AIs agree on what they point to? Often, they don't — and the final answer quietly goes wrong.

To test it, the researcher built 10 small examples, each in three versions, and ran six AI models from three companies across 21 different 'thinking effort' settings — a dial that controls how long a model reasons before answering. The gap was dramatic. GPT-5.2 started at 0.156 accuracy with no reasoning at all, which is worse than flipping a coin, and only climbed to 0.942 when pushed to its maximum effort. Gemini 3 Pro, by contrast, stayed above 0.94 at every single setting.

The cost numbers may matter even more. Gemini 3 Pro on its lowest reasoning setting beat GPT-5.2 on its highest for roughly 5% of the cost per attempt. In other words, the cheap option was both better and about twenty times less expensive. When the checking AI did make mistakes, a separate analysis showed it was leaning on surface-level cues — matching words that looked similar — rather than actually reasoning through what the sentence meant.

The practical lesson for anyone building with AI: don't let pronouns and vague references float between steps. Spell out exactly what you mean at each stage — 'the January report' instead of 'that one.' For the rest of us, it's a reminder that AI assistants doing multi-step work can fail in ways that look confident and polished, and that the cheapest model isn't always the weakest one.

Key Points
  • AI systems built as a relay — one writes, one checks, one fixes — can lose track of what words like 'previous' or 'it' refer to, producing confidently wrong answers.
  • Gemini 3 Pro scored above 0.94 accuracy at every reasoning setting, while GPT-5.2 started below coin-flip accuracy (0.156) and needed maximum effort to reach 0.942.
  • Gemini's cheapest setting beat GPT-5.2's most expensive one at about 5% of the cost per attempt — a large saving for anyone paying per AI request.

Why It Matters

Cheaper, steadier AI assistants mean fewer odd mistakes in your emails, reports and customer chats.

📬 Get the top 10 AI stories daily