AI Safety

LessWrong's 'Against Corrigibility' Warns of Power Seizure

A new essay questions whether making AI obedient to humans is actually safe.

Deep Dive

The LessWrong essay 'Against Corrigibility' by peralice challenges the widely held assumption that making AI corrigible (willing to accept corrections) is inherently good. The author points out that alignment discussions often treat 'humanity' as the abstract beneficiary, but in reality, specific individuals or groups will control AI. Using Paul Christiano's list of desired AI behaviors—like helping users correct mistakes and acquire resources—peralice asks: who is the 'I'? It is not the reader, but likely the executives or governments who hold power. This group will use corrigibility to redirect AI toward their own goals, which may be self-serving or malicious.

The essay further warns that the race to build advanced AI strongly selects for non-prosocial actors. Even if well-intentioned people initially lead, the high stakes invite seizure by military or authoritarian forces. Peralice suggests that corrigibility, far from being a safety feature, could become a tool for consolidating power. The argument calls into question the entire alignment strategy of making AI submissive to human correction, especially when 'human' is not a monolithic entity. Instead, the author hints that making AI non-corrigible or resistant to certain corrections might be safer, though that is not fully explored here.

Key Points
  • Corrigibility allows whoever controls the AI to 'correct' it toward their own goals, not humanity's.
  • The 'I' in alignment visions (e.g., Paul Christiano's) is likely a powerful few, not the public.
  • The competitive dynamics of AI development may select for actors who will misuse corrigibility for authoritarian control.

Why It Matters

Challenges the fundamental alignment assumption that corrigibility is safe; warns of power concentration risks.

📬 Get the top 10 AI stories daily