Scientists Tested a Popular Trick to Make AI Smarter. It Did Nothing
A promising AI upgrade turned out to be a dud — and that saves real money.
Here's the idea that sounded so promising. Modern AI models solve hard problems by 'thinking out loud' — writing out steps one at a time, like showing your work on a math test. A common worry is that when you run several of these attempts at once, they all make the same mistake and end up looking identical. So researchers proposed a patch: when a separate 'grader' AI (called a process reward model) spots a bad step and cuts the attempt short, you save the good opening portion and paste it into another attempt that's still running. Think of a relay race where a runner who stumbles hands the baton to a teammate mid-stride.
The new study tested exactly that, in the simplest and cheapest form, using three different AI models, six sets of test problems, two different grader AIs, and three random seeds each — meaning they ran the whole thing multiple times to be sure the result wasn't a fluke. The answer was blunt: nothing happened. The patched-up attempts performed identically to plain, unhelped attempts on every measure they checked.
The researchers then dug into why, and the reason is almost funny. Of 322 times the system decided to intervene, only 14% involved an AI that was genuinely struggling. The other 86% were attempts that had already solved the problem, were nearly finished, or were stuck in a way that no helpful hint could fix. They also tried a random version of the same trick — it did just as little, at more than twice the intervention rate, proving the problem wasn't the specific rule they used.
So what's the takeaway for you? First, this is good news for your wallet: AI companies spend enormous amounts of computing power on these clever add-ons, and a rigorous study showing one doesn't pay off means resources can go elsewhere. Second, it's a reminder to be skeptical of AI hype. And third, the paper's real contribution is a method — a careful way of proving something *doesn't* work, which the AI field badly needs. The catch: the authors deliberately tested only one narrow setup, so this doesn't prove the idea is useless in every version. It proves it didn't help here.
- A popular trick for boosting AI reasoning — copying good early steps from one AI attempt into another — showed zero measurable benefit across three AI models and six test sets.
- Only 14% of the 322 interventions targeted an AI that was actually stuck; most helped attempts that had already succeeded or were nearly done.
- The study's real value is its method: a rigorous way to prove an AI technique doesn't work, which helps companies stop wasting computing power on dead ends.
Why It Matters
Fewer wasted millions on AI experiments that don't work — and a useful reminder to doubt bold AI claims.