AI Safety

AI alignment field admits decade of missteps

A LessWrong retrospective accuses alignment researchers of accelerating unsafe AI for power

Deep Dive

LessWrong contributor Richard Ngo has published a five-part retrospective arguing that the AI alignment community has fundamentally derailed over the past decade. Initially framed as a hard scientific problem, Ngo claims the field pivoted toward incremental system improvements and power accumulation—often at the expense of safety-focused research. This shift, he argues, was driven by self-deceptive reasoning and fear, with alignment ideas directly contributing to the scaling of large language models and the development of ChatGPT.

The essay identifies four recurring mistakes the field has made: lobbying governments to take AGI seriously without clear strategic thinking, conflating alignment research with capabilities-maximizing work (e.g., building automated alignment researchers), over-trusting companies like Anthropic and OpenAI, and sacrificing analytical clarity for political conformity. Ngo traces these patterns back to a slippery slope logic—justifying harmful actions by assuming outcomes are inevitable—citing Sam Altman’s 2015 email to Elon Musk about founding OpenAI as a formative example.

Key Points
  • Richard Ngo (LessWrong) argues alignment researchers prioritized power and incremental gains over safety, accelerating unsafe AI development
  • Four major mistakes: government lobbying without clarity, conflating alignment with capabilities research, over-trusting Anthropic/OpenAI, and sacrificing intellectual integrity for conformity
  • The field’s ‘inevitability logic’ led to self-reinforcing acceleration, with alignment principles fueling LLM scaling despite ethical trade-offs

Why It Matters

Exposes how misaligned priorities in AI safety research may have made existential risks more likely by empowering unaccountable institutions.

📬 Get the top 10 AI stories daily