AI Safety

Byrnes' social-drive AGI alignment: human instincts as code

⚡Can human social instincts be coded into AGI to avoid misalignment?

Deep Dive

Steven Byrnes' latest alignment post tackles the unsolved problem of designing AGI motivations from scratch. His approach: start from human innate social instincts—like approval-seeking, norm enforcement, and pride—and find ways to implement them in a brain-like AGI. He argues that if humans can steer toward good outcomes, sufficiently human-like AGIs should be able to as well, provided they have prosocial drives. The post covers three failure modes: 1) the AGI might adopt the wrong moral circle (e.g., valuing only its creator's preferences), 2) norm-following may fail without a balance of power (AGI could override human norms), and 3) consequentialist reasoning could suppress social instincts over time.

Byrnes then examines how virtues emerge: through 'person-first' identification (copying admired individuals) or 'desire-first' (deep personal passion leading to self-image changes). He suggests the desire-first pathway is crucial for truth-seeking traits in AGI. The second half of the post shifts to backward reasoning from desiderata: technical constraints (code must be writable), strategic constraints (resilience against misaligned ASIs), and societal buy-in (the plan must sound reasonable). This is an early-stage, interconnected dump of ideas, seeking feedback and deconfusion.

Key Points
  • Byrnes proposes using human social instincts (approval, pride, norm enforcement) as inspiration for AGI motivation systems.
  • Three failure modes identified: wrong moral circles, power-imbalance breaking norms, and consequentialist desires overriding social drives.
  • Desire-first pathway (deep passion → self-image) is key for fostering truth-seeking virtues in brain-like AGI.

Why It Matters

Offers a novel alignment blueprint using human-like social drives, potentially safer than pure reward-maximization.

📬 Get the top 10 AI stories daily