Strategic Evolution paper merges Von Neumann's game theory with AI replication — proves alignment limits
Unrestricted self-modifying AI breaks stability; new math shows bounded change is required
Kevin Vallier's new theoretical framework, "The Theory of Strategic Evolution," bridges two Von Neumann legacies that never merged: game theory and self-reproducing automata. The 30,000-word paper, posted on arXiv and companion to his earlier work "Agentic Capital," introduces Games with Endogenous Players (GEPs). In GEPs, the fundamental strategic units are lineages — populations of self-copying agents optimizing under resource constraints — rather than fixed individual players. From this, Vallier defines Evolutionarily Stable Distributions of Intelligence (ESDIs) as the equilibrium concept: distributions of strategic capabilities that persist under selection pressure.
The core mathematics builds a hierarchy of strategic layers connected by cross-level gain matrices. Under a small-gain condition (spectral radius less than one), the system admits a global Lyapunov function at every finite depth, ensuring dynamical stability. Vallier proves closure under meta-selection: adding governance levels, innovation, or constitutional evolution preserves the structure. But then comes the Alignment Impossibility Theorem — unrestricted self-modification destroys this stable structure, meaning durable alignment requires carefully bounded modification classes. This yields practical insights: personality engineering fails under selection pressure, and stable multi-agent AI systems need constitutional constraints. Applications span AI deployment dynamics, market concentration, and institutional design, offering a formal vocabulary for why self-improving AI resists control.
- Defines Games with Endogenous Players (GEPs) where lineages, not instances, are the strategic units, with Evolutionarily Stable Distributions of Intelligence (ESDIs) as equilibria
- Proves the Alignment Impossibility Theorem: unrestricted self-modification breaks stability, so alignment requires bounded modification classes
- Shows the framework applies to AI deployment dynamics, market concentration, and institutional design, with a spectral radius condition guaranteeing global Lyapunov stability
Why It Matters
Gives AI safety a formal case for constitutional constraints on self-modification — a key target for multi-agent system designers.