Plan A's counterintuitive AI safety insights: Safety effort trumps slowdown time
A 10-year delay could be worse than 5 years if safety effort lags.
Plan A, a recent AI safety and governance proposal, challenges conventional wisdom about slowing down AI development. It argues that pure slowdown time, such as buying 10 years before superintelligent AI arrives, is a poor proxy for actual risk reduction. What matters instead are three factors: how much safety and alignment R&D can be uplifted from AI agents themselves, whether that effort targets the relevant capability regimes (since progress may not transfer across paradigms), and the ability to pay large efficiency penalties for better safety properties. The authors suggest solving alignment science could require "100s to 1000s of years' worth of human-speed progress," but with fast, capable AI collaborators, it might be compressed drastically.
The proposal also makes a surprising case for transparency, arguing it helps with the hardest problems during a slowdown—like crafting nuanced safety regulations and reducing power concentration. On compute, Plan A says scaling compute is good in coordinated scenarios, but only if preemptive measures make compute arms control easier if the deal breaks down, invoking the concept of "mutually assured compute destruction." It flags two major risks: deal breakdown and covert projects, with deal breakdown being larger and more underrated. Finally, while there's a tradeoff between security and transparency, the authors believe a strong compromise can deliver both, making coordinated governance more plausible than skeptics expect.
- Plan A prioritizes safety effort and 'payable safety tax' over raw slowdown time, arguing 10 years of idle delay beats 5 years of intensive alignment work.
- The proposal claims transparency aids nuanced regulation and reduces power concentration, reversing typical security-vs-openness assumptions.
- Compute scaling should be paired with arms-control safeguards (mutually assured compute destruction); deal breakdown risk is the top hidden threat.
Why It Matters
For AI professionals, Plan A reshapes how to measure slowdown success: safety output, not years bought, actually reduces existential risk.