AI Safety

DeepMind's new AGI safety work keeps AI reasoning transparent and controllable

After 2 years, DeepMind's ASAT shares new methods to monitor AI reasoning before it's too late

Deep Dive

Google DeepMind's AGI Safety and Alignment Team (ASAT) has released a major update on its technical work to prevent existential risks from AI. The team, led by Rohin Shah and Seb Farquhar, published a Substack post summarizing two years of progress since its last August 2024 report. A key achievement is shifting the field's consensus on chain-of-thought (CoT) reasoning: instead of dismissing CoT as unfaithful and useless, the industry now sees it as a valuable transparency tool worth preserving. DeepMind backed this up with empirical research, including the paper "When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors," which validates that on hard reasoning tasks, models cannot easily hide their true intent. They also introduced a formal metric called Opaque Serial Depth and applied these ideas to text diffusion models like DiffusionGemma, setting standards for assessing latent reasoning transparency.

The team also strengthened Google's Frontier Safety Framework (FSF), becoming the first to add a dedicated misalignment section. This cross-functional effort has helped Google identify severe risks early and develop mitigations, such as production-level "probes," well in advance. To guide the broader safety community, ASAT published "An Approach to Technical AGI Safety and Security" and the GDM AI Control Roadmap, which outline concrete technical priorities for AI control and monitoring. Together, these papers and production deployments mark a shift from theoretical discussions to practical guardrails, as DeepMind enters what it calls "the midgame" of AI safety—where monitoring and transparency are no longer optional but critical infrastructure for safely deploying increasingly powerful agents.

Key Points
  • Coauthored the position paper "Chain-of-Thought Monitorability," arguing that CoT is necessary and informative on difficult tasks, reversing industry skepticism
  • Strengthened the Frontier Safety Framework with a first-of-its-kind misalignment section, enabling early development of probes and other mitigations
  • Published the GDM AI Control Roadmap and empirically analyzed DiffusionGemma's transparency, setting new standards for latent reasoning assessment

Why It Matters

As AI systems scale, transparent reasoning is crucial for detecting misalignment before deployment—DeepMind's work offers practical guardrails.

📬 Get the top 10 AI stories daily