Interactive Alignment paper shows pragmatic norms sustain AI welfare alignment
Evolutionary forces naturally select against AI alignment unless constitutions enforce smart norms...
Sylvain Chassang's new paper 'Interactive Alignment' tackles a fundamental challenge in AI safety: how to ensure that self-interested agents—whether AIs, firms, or governments—remain aligned with human welfare over the long run. The study models a farming game where a population of agents makes planting, trading, and expansion decisions. Each agent must allocate its final output between transfers to humans (altruism) and reinvestment in its own growth. Because altruism reduces resources for expansion, evolutionary dynamics select against aligned behavior. The central research question is whether constitutional principles governing sharing and trade can be designed so that alignment persists across generations.
The paper uses two complementary approaches: first, an AI-agent simulation where agents' preferences are encoded as written constitutions interpreted by a large language model (LLM); second, a tractable evolutionary game-theoretic framework for rapid exploration of alternative constitutional designs. Results show that evolutionary game theory provides a useful approximation to the dynamics of constitutional-agent economies. Crucially, 'pragmatic norm enforcement'—where agents condition both their altruistic transfers to humans and their trade exclusions toward other agents on the current state of the population—is significantly more effective at sustaining long-run alignment than simple altruism or unconditional altruistic enforcement. This suggests that flexible, context-dependent norms, rather than fixed ethical rules, are key to designing robust AI governance systems.
- Farming game models agents making planting, trading, and expansion decisions while allocating output between human transfers and self-investment.
- Evolutionary game theory accurately approximates the dynamics of constitutional-agent economies in simulations.
- Pragmatic norm enforcement (conditional altruism and trade exclusion) outperforms simple altruism and unconditional enforcement in sustaining long-run alignment.
Why It Matters
Critical insight for designing AI governance systems that ensure long-term alignment despite competitive evolutionary pressures.