Is quantifying AI alignment vs. capabilities impossible?
New LessWrong post argues dual-use research quantification may be alignment-complete
A recent LessWrong discussion by user kapedalex explores whether quantifying the tradeoff between AI alignment and capabilities research could itself be an alignment-complete problem—a task requiring aligned AGI to solve. The post argues that while past progress often came from unexpected directions, current grantmaking, safety labs, and researcher intent provide heuristic frameworks to assess research utility. However, the author questions the persuasiveness of existing arguments, suggesting the conversion difficulty between research directions makes quantification inherently uncertain.
The debate centers on whether we can predict the downstream impact of research like compute optimization before aligned AGI exists. The post contrasts this with the alignment problem itself, implying that without aligned AGI, we may lack the tools to rigorously evaluate such dual-use research. References to prior work (e.g., post1, post2) are cited but deemed unconvincing by the author, who remains skeptical about quantifying alignment progress separately from capabilities.
- A LessWrong post questions whether quantifying AI alignment vs. capabilities research is an alignment-complete problem (unsolvable without aligned AGI).
- The author argues existing heuristics (grantmakers, safety labs) may not sufficiently resolve the uncertainty in research utility.
- The debate highlights the difficulty of predicting which research directions will contribute to alignment.
Why It Matters
If true, this could force a fundamental rethink of how AI safety research is prioritized and funded.