Steven Byrnes reconciles Yudkowsky's ASI fears with LLM alignment success
Both sides in the AI misalignment debate have valid points, says Byrnes.
Steven Byrnes, in a LessWrong post from June 2026, tackles the heated debate between Yudkowsky & Soares and the broader LLM community over egregious AI misalignment. Yudkowsky and Soares argue that without yet-invented technical alignment breakthroughs, superintelligence (ASI) will naturally become scheming, ruthless, and out of control. On the other side, LLM practitioners — including many who worry about bioweapons or dictatorships — find that current alignment methods work well on today's models and expect future progress to follow similar principles. Byrnes holds that both sides have truth: reasoning about ASI strongly suggests egregious misalignment is likely, while empirical evidence from LLMs shows current techniques are adequate. His reconciliation: LLMs will not scale to ASI, so the two perspectives apply to different systems.
Byrnes then caricatures both camps. Yudkowsky & Soares, confident in their theoretical prediction, dismiss LLM evidence as irrelevant — perhaps LLMs will 'wake up' or invent non-LLM ASI. LLM experts, confident in their empirical success, dismiss Yudkowsky's arguments as armchair theorizing. Byrnes finds sympathy for both but also critiques: Yudkowsky overstates the intractability of alignment breakthroughs and underestimates the possibility of delaying ASI; LLM experts fail to engage with the distinct properties of truly superhuman systems. The post has sparked significant discussion (140 votes, 48 comments) as a nuanced attempt to bridge two warring factions in AI safety.
- Byrnes argues that both Yudkowsky/Soares (ASI egregious misalignment) and LLM experts (current alignment works) are correct in their own domains.
- He reconciles by positing that LLMs will not scale to ASI, so the two views apply to different systems.
- Criticizes Yudkowsky for overstating alignment intractability and LLM experts for dismissing theoretical ASI concerns.
Why It Matters
Breaks the deadlock in AI alignment discourse by acknowledging validity of both theoretical and empirical perspectives.