New AI Trick Lets Assistants Learn From Mistakes They Never Made
This could make AI helpers cheaper, faster and less likely to repeat errors.
Today's AI helpers — the kind that can actually do things for you, like book a flight or fix code — mostly learn from what really happened. If a step fails, they remember the error. But they never stop to ask "what if I'd tried something else?" That's like a person who only learns from mistakes they already made, never from ones they narrowly avoided.
A new research system called COUNTERMEM fixes that. After a failed action, it makes a copy of the situation — think of saving your game right before a risky jump — then tries a few nearby alternatives on that copy. A checker decides which alternative actually worked. That checker can be a set of software tests, a math proof checker, or a puzzle solver. Whatever worked gets saved alongside the original mistake, along with the conditions where that fix applies. A separate learned policy then decides whether to use the saved fix or skip it, to avoid wasting time.
The results are promising. Across 12 test settings in six areas, the system improved two popular AI-agent methods — ReAct and Reflexion — by about 12.6 percentage points on average. It also cut the number of tokens used by 7.7% to 42%. Tokens are the metered units that make AI expensive, so fewer of them means lower bills. The researchers also found that if you remove the verification step or the saved memory, the benefits shrink. Worse, applying a fix to a situation it doesn't fit can make the AI do worse than before.
The honest catch: this is a research paper heading to a 2027 conference, so the code isn't public yet. And it needs a reliable way to check answers. That's easy for code and math, where correctness is objective, but much harder for messy real-world tasks like negotiating a refund or planning a trip. The savings also exclude the upfront cost of training the system.
- AI helpers that replay 'what if' scenarios succeeded about 12.6% more often in tests
- The approach used up to 42% fewer tokens — the metered units that drive AI costs
- It only works where answers can be checked automatically, like code tests or math proofs
Why It Matters
Cheaper, more reliable AI helpers could take on more everyday chores — but only where results can be verified.