AI Safety

LLMs lack 'strong generalization' – the hidden human superpower needed for AGI

Why do AI experts outperform novices with the same LLM? The answer may redefine intelligence.

Deep Dive

Stuart_Armstrong, writing on the AI Alignment Forum in July 2026, identifies a critical gap in LLMs: 'strong generalization.' He defines it as the set of abilities—situational awareness, out-of-distribution generalization, symbol grounding, adaptive world modeling, long-range planning, and anomaly detection—that let humans work and plan in novel situations while pursuing objectives. Humans use this semi-instinctive process without conscious awareness, making it hard to detect its absence in AI. Armstrong argues that LLMs lack strong generalization entirely; they are 'vastly knowledgeable and articulate' but cannot infer beyond their training data. This explains five mysteries: why subject-matter experts get much more from LLMs than amateurs, why LLMs show spiky abilities with obvious failures, why they need enormous datasets and constant retraining, why they seem perpetually on the verge of AGI but never arrive, and crucially, why they succeed on long-horizon benchmarks but cannot transpose that performance to the real world.

For professionals, Armstrong's theory reframes the AGI debate: the bitter lesson of scaling may not suffice because the missing ability is qualitative, not quantitative. LLMs remix data non-trivially, but human engineers provide the strong generalization by keeping them on track. This suggests that current architectures cannot achieve true general intelligence without a fundamental addition—one we don't yet know how to code because we don't understand our own cognitive processes. The implication is stark: without strong generalization, LLMs will remain brittle, failing not at the moment of error but because they were attempting an impossible task from the start. The benchmarks that purport to teach long-term planning may actually be gamed without generating real-world transfer. Armstrong doesn't offer a fix, but he issues a clear warning: 'we don't know how to code it up, and we don't recognize its absence or presence in other entities.' For AI researchers and builders, this is a call to look beyond scaling and benchmark performance toward the elusive, human-like ability to truly generalize.

Key Points
  • LLMs succeed on long-horizon benchmarks but fail to transfer learning to novel real-world tasks, indicating a lack of true generalization.
  • Subject-matter experts effectively supply the missing 'strong generalization' by using their own understanding to keep the LLM on track, explaining the performance gap between experts and amateurs.
  • Humans cannot consciously perceive their own strong generalization, so they attribute LLM failures to specific bugs rather than recognizing a fundamental architectural missing piece.

Why It Matters

This theory explains the current AGI plateau and warns that scaling alone won't bridge the gap—a fundamentally new capability is needed.

📬 Get the top 10 AI stories daily