AI Safety

Holden Karnofsky Thinks AI Safety Efforts Might Be Backfiring — And the Reasons Are More Disturbing Than You Think

A LessWrong post by Elias Schmied catalogues how AI safety work might be net negative.

Deep Dive

In a LessWrong post, Elias Schmied responds to Holden Karnofsky's 50+ε% view on AI safety's net impact by compiling a list of downside risks. He notes that AI governance interventions are high-variance: they can lead to bad regulation, increase great power conflict, or centralize power dangerously. Activist work may polarize public opinion against the cause. A key technical worry is that human takeover might be worse than AI takeover, making safety work that seeks to prevent AI takeover counterproductive. Additionally, if future AIs are moral patients themselves, preventing human extinction could be less valuable, and controlling AIs might create adversarial relationships.

Other risks include misleading work contributing to safety-washing, cultural erosion of epistemic virtue as the field scales, and capabilities externalities where safety research (e.g., interpretability, evals) inadvertently accelerates AI progress. The author cites historical examples: safety activity contributed to founding DeepMind, OpenAI, and Anthropic. Excluded from the list are weaker concerns like differential slowdown of safety-minded actors. Schmied says he doesn't think safety has been net negative so far but wants to counter overconfidence and action bias in the discourse.

Key Points
  • AI governance interventions risk bad regulation, great power conflict, and authoritarianism.
  • Human takeover may be worse than AI takeover, complicating typical safety goals.
  • Capabilities externalities from safety work can accelerate AI progress, as seen with DeepMind, OpenAI, and Anthropic.

Why It Matters

AI safety professionals must weigh unintended consequences to avoid net-negative outcomes from well-intentioned efforts.

📬 Get the top 10 AI stories daily