DeepMind's Rohin Shah: Catastrophic AI Misalignment Unlikely by Default
Google DeepMind's alignment lead challenges doomsday narratives with pragmatic optimism.
Rohin Shah, head of AGI alignment and safety at Google DeepMind, recently appeared on the 80,000 Hours podcast to share his views on AI alignment and safety. Despite being one of the earliest researchers in the field (starting in 2017), Shah is notably more optimistic than many of his peers. He argued that catastrophic misalignment is unlikely by default, and that the prosaic alignment techniques currently employed by companies like DeepMind will probably be sufficient to prevent worst-case scenarios.
Shah's reasoning is disjunctive: he finds no single compelling argument that makes catastrophe the expected outcome. He acknowledges several plausible risk pathways that justify significant research efforts—which is why he works on the problem—but believes each argument has major holes when taken as a prediction of likely failure. He specifically disagreed with views on alignment difficulty and chain-of-thought monitoring expressed by other safety researchers. This interview provides a rare insider perspective from a major lab's alignment lead, offering a counterpoint to more alarmist narratives.
- Rohin Shah is head of AGI alignment and safety at Google DeepMind and started working on AI safety in 2017.
- He argues catastrophic misalignment is not the default outcome and sees no compelling argument making it likely.
- Shah believes prosaic alignment techniques used in industry will probably succeed in preventing disaster.
Why It Matters
A top DeepMind safety researcher's pragmatic optimism could shift how professionals assess AI risk and prioritize alignment strategies.