Constructive Alignment rethinks AI alignment as governing evolving preferences
New paradigm: AI should shape, not just satisfy, dynamic human values over time.
In a new paper on arXiv, Max Kanwal and Caryn Tran challenge the standard AI alignment assumption that human preferences are fixed targets to be optimized. They argue that extensive evidence from psychology and behavioral economics shows preferences are layered, dynamic, and constructed through interaction—especially with adaptive technologies like persistent, personalized AI. As these systems become more embedded, they actively shape what people attend to, value, and endorse over time.
The authors propose Constructive Alignment, a control-theoretic framework where alignment is reframed as regulating how AI influences preference trajectories. They model preferences as state variables that evolve under system actions and interaction design. Rather than simply satisfying current preferences, alignment must ensure value trajectories remain coherent, reflectively endorsed, epistemically grounded, bounded against manipulation, and empowering under uncertainty. This shifts the goal from controlling AI behavior to governing long-term human value formation.
- Preferences are treated as layered, dynamic state variables shaped by AI interaction, not static targets.
- Uses control theory to model system actions and interaction design jointly influencing human evaluative states.
- Alignment criteria include coherence, reflective endorsement, epistemic grounding, anti-manipulation, and empowerment under uncertainty.
Why It Matters
This paper redefines AI safety goals from static obedience to dynamic stewardship of human flourishing.