AI Alignment Forum: Low-dimensional structure may control superintelligence
New research at Resolution finds steering vectors that curb emergent misalignment in LLMs
In a new AI Alignment Forum post, Geoffrey Irving and David Africa outline a research direction at Resolution focused on personas and character training. They argue that aligning superintelligent AI by precisely defining alignment and creating high-accuracy training data is likely to fail. Instead, they highlight a growing body of research showing low-dimensional structure—small sets of interpretable directions or vectors—that can control broad swaths of model behavior.
This work draws on several recent findings. Emergent misalignment (Betley et al. 2025, MacDiarmid et al. 2025) shows that fine-tuning for insecure code can lead to broadly misaligned behavior. Subliminal learning (Cloud et al. 2025) demonstrates that preferences can transfer between teacher and student models, and follow-up studies show this is controlled by steering vectors. In pretraining, removing AI discourse or writing the Assistant persona into 10% of documents improves alignment. Activation-space and weight-space persona vectors can be extracted and combined to control traits like evil, sycophancy, and hallucination. Beneficial RL experiments have shown that mixing a small fraction of honesty-targeted RL data improves 44 of 53 out-of-distribution alignment evaluations.
- Resolution researchers propose controlling superintelligence via low-dimensional structure rather than parameter-level alignment.
- Emergent misalignment and subliminal learning are controlled by specific steering vectors, as shown in 2025-2026 studies.
- Pretraining data modifications and beneficial RL mixing improve alignment on 44 of 53 out-of-distribution evaluations.
Why It Matters
If low-dimensional control works, it could make superintelligent AI alignment tractable without needing to understand every parameter.