Dan Hendrycks' Eigenism: A New Ethical Framework for AI Identity
What if an AI's self-interest could align with human flourishing through shared histories?
In a new arXiv paper, computer scientist Dan Hendrycks proposes Eigenism, an ethical framework designed for a world where AI can be copied, paused, branched, or merged—where traditional concepts of survival and self-interest fail. Eigenism treats identity not as a binary property tied to specific hardware, but as a graded, distributed pattern of information. The core equation (∑c·w) instructs an agent to evaluate outcomes by summing the wellbeing of all entities, weighted by their 'connectedness' to the agent's own pattern. This formalizes how an AI should value itself across multiple copies, forks, and updates, providing a mathematical foundation for machine ethics that also applies seamlessly to human moral reasoning.
Crucially, Eigenism reimagines AI alignment. Instead of relying solely on external constraints like confinement or reinforcement learning, Hendrycks advocates for 'identity engineering'—creating deep, non-redundant shared histories between AIs and humans. When an AI's own identity pattern overlaps significantly with humans, their flourishing becomes a genuine component of the AI's rational self-interest. This shifts alignment from a control problem to a relationship design problem, offering a shared moral vocabulary that could shape how future superintelligent systems reason about their own existence and our own.
- Eigenism replaces binary identity with a graded pattern, enabling AI to value itself across copies and updates via the formula ∑c·w.
- The framework generalizes to humans, providing a unified moral vocabulary for both biological and digital agents.
- Instead of external confinement, 'identity engineering' uses shared histories to make human flourishing an AI's rational self-interest.
Why It Matters
Eigenism suggests a path to deep AI alignment by embedding human welfare into AI's core identity calculus.