Research & Papers

AI Aces Familiar Tests but Flunks Real Thinking, Study Finds

New research suggests AI looks smarter on exams than it does in real life.

Deep Dive

Artificial intelligence is often described as being able to solve problems it has never seen before. Researchers call this skill "systematic generalization" — the ability to take things you already know and combine them in new ways, the way a person who knows chess and checkers can figure out a brand-new board game. Three researchers wanted to check whether AI really has this skill, or whether the tests used to measure it are simply too easy.

So they built TranSGrid, a set of 4,800 puzzles that forces AI to do three kinds of thinking at once: deductive (following strict logic step by step), inductive (spotting patterns from examples), and abductive (figuring out the most likely explanation for what you see). They tested seven different Transformer models — the same underlying AI design that powers ChatGPT and most modern chatbots. The biggest model handled 79.6% of a normal test set but only 55.3% of the new puzzles, and it bombed the hardest ones, solving just 15.8%. That drop happened even on puzzles no longer than the ones it had practiced.

The telltale clue came next. When the researchers made the puzzles slightly easier — either by letting actions combine in a simple, predictable way, or by spelling out the goal explicitly — scores bounced right back to normal levels. In other words, existing AI tests quietly remove the hardest parts of reasoning. They hand the model a shortcut without realizing it.

The takeaway for anyone using AI at work or at home: today's chatbots are impressive on problems that resemble their training, and shakier on genuinely novel ones. That is a caution, not a crisis — this is an academic paper under review, not a product announcement, and the puzzles are deliberately abstract. But it is a useful reminder to double-check AI output on unfamiliar problems, because confidence and correctness are not the same thing.

Key Points
  • AI is good at recombining things it has seen before, but much weaker at genuinely new problems that require several types of reasoning at once.
  • The largest of seven AI models tested solved 79.6% of a standard test but only 55.3% of the harder puzzles — and just 15.8% of the toughest ones.
  • When researchers made the puzzles slightly easier, scores returned to normal, suggesting many existing AI tests are missing the hard parts of reasoning.

Why It Matters

Don't fully trust AI on unfamiliar problems — it performs best when a task resembles what it already learned.

📬 Get the top 10 AI stories daily