New AI Solves Puzzle Tests by Showing Its Work
An AI that shows its work could make the tech you rely on more trustworthy.
A researcher named Deblina Kar published a new approach to one of AI's toughest challenges: solving puzzles that require real flexible thinking. The test, called ARC (a set of colored-grid puzzles, like an IQ test for machines), is famously hard for computers. Most AI systems fail badly. This new framework breaks each puzzle into stages — first spotting simple rules about shapes and colors, then combining patterns, then inferring bigger structural relationships. If one stage fails, the next picks up the work, reusing what came before. The result: more than 95 percent accuracy, solving 230 of 240 of the hardest ARC-AGI-2 tasks.
The real story isn't the score — it's how it gets there. Most powerful AI today is a black box: you ask, it answers, and nobody can explain why. This system leaves a readable trail of reasoning, like a student showing their math work instead of just writing down a guess. That matters because when AI weighs in on your loan application, your medical scan, or your job resume, "trust me" isn't good enough. A trail you can inspect means mistakes can be found, argued with, and fixed.
The catch: this is a research paper, not a product. You can't download it, and grid puzzles are a long way from the messy real world of tax documents, conversations, and hospital records. The approach also leans on hand-crafted reasoning rules, which may not stretch to every kind of problem. And one author's results, however impressive, need other researchers to repeat them before anyone calls it a breakthrough.
Still, the direction is telling. If AI can explain itself, regulators can audit it, companies can defend their choices, and you can push back when something seems wrong. Expect the next few years to focus less on raw intelligence and more on whether we can trust — and check — what machines tell us.
- It solved 230 of 240 of the hardest visual reasoning puzzles, scoring above 95 percent overall.
- Instead of one giant black box, it chains simple rules together and leaves a readable reasoning trail.
- It's still a research paper, not an app — and grid puzzles are far simpler than real life.
Why It Matters
AI you can question could be trusted with hiring, loans, and medical calls — where blind answers aren't enough.