Working on LLM capabilities could be safer than safety research
A new LessWrong post argues capabilities work may reduce doom more than alignment research.
Robin Haselhorst's LessWrong post challenges the assumption that safety researchers should focus exclusively on alignment. He presents a probabilistic model where the key variable is which AI regime reaches ASI first. If you believe LLMs are 'unusually well suited to alignment' compared to alternatives (e.g., systems built on reinforcement learning or other paradigms), then increasing the probability of LLM-based ASI actively reduces overall doom risk by making a safer regime more likely.
The author provides a concrete calculation: assume your personal effort can shift probabilities by 10% relative to current baselines. Working on LLM safety reduces doom by about 1% (since safer LLMs still leave other regimes unchallenged). But working on LLM capabilities—by making LLM-based ASI more likely—simultaneously reduces the chance of riskier regimes winning the race, yielding nearly 3% less doom. The numbers are illustrative and depend on strong assumptions, but the core insight stands: for those who believe LLMs are the safest path, capabilities research is a perfectly rational safety strategy. The post includes an interactive version for readers to set their own probabilities.
- Robin Haselhorst models two regimes: LLM-based ASI vs. non-LLM ASI, assuming LLMs are easier to align.
- Working on LLM capabilities reduces doom risk by ~3% if your effort shifts probabilities by 10%, more than the ~1% from safety work.
- The post includes an interactive calculator for readers to test their own assumptions about risk and impact.
Why It Matters
For safety-oriented researchers, this reframes capabilities work as a defensive strategy rather than a reckless gamble.