Do AI Assistants Actually Learn From Experience? New Test Says Not Always
Before you trust an AI with bigger jobs, will it actually get smarter with practice?
Most of us assume that AI assistants learn as they go. You correct it once, and it remembers for next time. But a new research test called AhaBench asks a more direct question: when an AI gets useful information, does it actually carry that skill into later, harder versions of the same problem? On this new test, the honest answer is: not always — and even top models are inconsistent.
The researchers from Princeton and other institutions created three scenarios. In Aha-Puzzle, an AI sees hints to solve hidden-state puzzles, then must solve similar puzzles without hints. Aha-Euler teaches a model math problems and requires it to solve related new problems. Aha-Vending puts a simulated agent in charge of vending machines, handling repairs and restocking while dealing with delayed feedback on profits. All these test whether an AI improves based on prior experience.
The results show a split personality. Claude Opus 4.6 scored highest overall in later-performance score (64.3) and displayed the most improvement from its starting point. But models that use support well are not necessarily those that learn most over time. In some math tasks, models reached 78-100% accuracy when fully taught, but dropped to as low as 0% when only given answers. Similarly, puzzle agents could solve problems with cues, yet fail when cues were removed.
Why does this matter? Companies are betting on AI that can handle long-term responsibilities — like running customer accounts or monitoring store operations. If an AI handles a task perfectly in a conversation, that's not proof it will handle next week's version faster or better. The benchmark offers separate scores for starting skill, skill after experience, and measured improvement, so businesses can see whether an AI truly grows on the job or just performs in the moment.
- A new test called AhaBench checks whether AI models actually improve on long-term tasks after gaining experience.
- Claude Opus 4.6 was the best performer, improving by 25.8 points, but models that used help well didn't always transfer that learning to new situations.
- The test includes three real-world scenarios: puzzle solving, math learning, and managing vending machines with delayed feedback.
Why It Matters
If AI can't learn from past experience, businesses and consumers can't rely on it for long-term tasks.