AI Safety

Alignment work more promising than control, says Alec Harris

Control's narrow window vs. alignment's scalable future with AI aid.

Deep Dive

Alec Harris challenges the prevailing emphasis on AI control research by arguing that alignment work offers a more promising path. He identifies the "control window" as the period where frontier AIs, which could otherwise take over, are kept in check by safety protocols. This window, he contends, is narrow: even with strong control measures (e.g., untrusted monitoring, honeypotting), the required intelligence threshold to prevent takeover increases only modestly. The marginal gains from additional control research shrink quickly as models become smarter, making control a limited lever.

Harris then applies a similar framework to alignment. Most alignment interventions also do not scale to ASI, but they benefit from a forward-chaining theory of change: moderately aligned AIs can help produce alignment work for more capable successors—what labs already rely on. The alignment window, he argues, is both wider and more "right-centered" (peaking when AI labor is highest). He calculates that shifting effort to an 8:1 ratio (alignment to control) would maximize safety gains, as alignment can compound returns while control hits diminishing returns. The article urges the community to reconsider resource allocation before the window narrows further.

Key Points
  • Control window starts at ~200 intelligence points and widens only ~25 points even with strong protocols, offering diminishing returns.
  • Alignment leverages a forward-chaining loop: aligned AIs help align smarter models, enabling scalable safety work.
  • Harris proposes an 8:1 effort ratio (alignment:control) to maximize long-term safety against ASI threats.

Why It Matters

Shifts AI safety debate from short-term control to scalable alignment, urging rebalancing of research priorities.

📬 Get the top 10 AI stories daily