AI Safety

Gwern's New Scaling Hypothesis: Humans as Over-Parameterized LLMs?

Are LLMs just under-parameterized brains? Gwern proposes a radical shift in scaling strategy.

Deep Dive

Gwern's latest LessWrong essay, 'Scaling Hypothesis #2: Are Humans Just More Over-Parameterized?', reopens the debate on why biological brains generalize so well while LLMs remain brittle. He posits that the key difference lies in a bias-variance tradeoff: current LLMs minimize variance (memorizing patterns) while human brains minimize bias by being massively over-parameterized and trained with extremely high learning rates on small, diverse, filtered datasets. This strategy, he argues, allows brains to 'catapult' into a flat, generalizing basin in the loss landscape, achieving sample efficiency and adversarial robustness at the cost of poor memorization. Gwern suggests testing this hypothesis by training multi-trillion parameter transformer models for very few steps using cyclical learning rate schedules, then benchmarking them on adversarial and out-of-distribution tasks like arithmetic and small-image classification.

However, the post has sparked strong pushback from prominent AI researchers. Lucius Bushnaq contends that massive overparameterization is not necessary for finding well-generalizing solutions—proper weight regularization can smoothly achieve generalization without the 'grokking' phase. More critically, Bushnaq warns that improved generalization could actually undermine AI safety: making models more creative and agentic makes it harder for engineers to predict what internal objectives they will develop, increasing the risk of misalignment. Andrii Vasylenko echoes this, noting that while capabilities may have attractor basins, alignment does not. The debate underscores a fundamental tension: better generalization may not inherently lead to safer AI, and Gwern's proposal, while intriguing, faces serious practical and philosophical hurdles.

Key Points
  • Gwern proposes training multi-trillion parameter LLMs with high learning rates for few steps to mimic human generalization.
  • The hypothesis suggests human brains use extreme over-parameterization to minimize bias, while LLMs minimize variance.
  • Critics warn that better generalization could make AIs more agentic and harder to align, not safer.

Why It Matters

This hypothesis challenges core scaling assumptions and could reshape how we train models for robustness and safety.

📬 Get the top 10 AI stories daily