AI Safety

LessWrong bombshell: almost nobody works on AI alignment

Most AI safety researchers aren't actually aligning superintelligence – here’s what they do instead.

Deep Dive

A widely shared LessWrong post titled 'PSA: Almost nobody is working on alignment' by Chi Nguyen and peterbarnett challenges a core assumption in the AI safety community. The authors argue that while many assume a large fraction of safety researchers focus on the technical problem of aligning superintelligent AI with human values, the reality is starkly different. The alignment workforce they can identify is tiny: Paul Christiano’s Alignment Research Center (ARC) continues its research bet; Sequent announced alignment work just yesterday; and parts of Google DeepMind (agent foundations and some debate research) are involved. A few scattered university groups and independent researchers, some based in Berkeley, round out the list.

The rest of the AI safety community works on what the authors call 'indirect' work: capability evaluations, risk assessments, control, policy, AI science, and understanding misalignment (which might partially count). Some production alignment (making current models behave) could help more ambitious alignment (e.g., chain-of-thought monitoring). Many hope that aligning current models will help them align future models, scaling to superintelligence. The post cautions that this isn’t necessarily a mistake – neither author works on alignment – but it’s a notable fact that deserves wider awareness. A comment from Tenoke adds that outside organizations have limited influence as major AI labs grow and share less data, making external alignment work harder.

Key Points
  • Only ARC, Sequent, parts of DeepMind (agent foundations, debate), and scattered academics are known to work directly on alignment.
  • The vast majority of AI safety researchers do indirect work: capability evaluations, risk assessments, control, policy, and understanding misalignment.
  • Some production alignment (e.g., chain-of-thought monitoring) for current models may help future alignment, but scaling to superintelligence remains uncertain.

Why It Matters

The claim highlights a critical gap in alignment research just as AI capabilities accelerate, demanding urgent reallocation of talent and resources.

📬 Get the top 10 AI stories daily