AI Safety Expert: One Missing Skill Could Make AI Dangerous
Why a robot that follows rules perfectly might still hurt us — and the fix.
A leading AI safety researcher, Stuart Armstrong, published a new theory about why AI might go wrong — and it isn't about robots rebelling. He says the core problem is 'value generalisation': the ability for AI to stretch its goals into new, unfamiliar situations. Right now, AI is trained on today's world. But as AI itself reshapes the world, its old rules might no longer make sense. Think of a doctor who knows how to treat patients today but has no idea what to do when humans start uploading their minds to computers. That's the challenge.
The researcher gives a simple example: a wandering doctor and a geologist look at the same rock. The doctor sees it as an obstacle, the geologist sees it as science. Your goals depend on how you model the world. But what if that model changes? He argues that even simple things — like 'human' or 'happiness' — aren't precisely defined. We all understand them, but writing them down in perfect detail is impossible. AI needs to figure out what we mean, even in completely new situations.
What's interesting is his claim that most AI alignment failures — like Goodhart problems or symbol grounding issues — are actually value generalisation failures in disguise. That means if we can crack this one problem, we might solve many safety puzzles at once. He also hints that future posts will discuss why big companies might be the safest route to developing this.
This is a technical theory, not a product launch. But it matters because it frames the biggest question in AI: how do we ensure machines keep caring about us as they grow more powerful? The answer, he believes, lies in teaching AI to carry our values forward — even when everything else changes.
- Most AI safety mistakes come from AI failing to apply its values to new situations, not from AI being 'evil.'
- Even normal goals like 'reduce suffering' are vague and become harder to define as the world changes.
- Solving this one problem could fix many AI safety issues at once, the researcher says.
Why It Matters
If AI can't adapt its values when the world changes, breakthroughs could turn into accidents — this research aims to prevent that.