AI Safety

Stuart Armstrong argues LLMs lack 'strong generalisation,' explaining the AGI gap

A hidden human superpower that LLMs cannot replicate, despite data scale.

Deep Dive

Stuart Armstrong argues that the persistent gap between large language models and true AGI stems from a missing ability he terms 'strong generalisation.' This is not a single skill but a cluster of interconnected capacities: situational awareness, out-of-distribution generalisation, symbol grounding, adaptive world modelling, long-range planning, hierarchical planning, anomaly detection, and more. Humans exercise these abilities semi-instinctively, often without conscious awareness, making them hard to define or code. LLMs, by contrast, are powerful remixers of training data but cannot truly generalise beyond it.

Armstrong points to five mysteries that strong generalisation solves: why subject matter experts extract more value from LLMs than amateurs; why LLMs show a bizarre mix of brilliance and obvious stupidity; why they require constant retraining on huge datasets; why they always seem on the verge of AGI but never arrive; and why they can ace long-horizon benchmarks without transferring those skills to real-world planning. The bitter lesson and scaling arguments fail because LLMs saturate benchmark metrics without acquiring genuine generalisation. Until strong generalisation is understood and replicated, true AGI remains elusive.

Key Points
  • Strong generalisation is a mix of ~8 human abilities including situational awareness, symbol grounding, and hierarchical planning that LLMs lack.
  • LLMs succeed on long-horizon benchmarks without learning actual planning or generalisation, debunking scaling arguments.
  • Subject matter experts compensate for LLM weaknesses by providing the strong generalisation missing in the models.

Why It Matters

This analysis challenges the scaling orthodoxy, suggesting that more data alone won't bridge the AGI gap.

📬 Get the top 10 AI stories daily