AI Safety

Byrnes' social-drive AGI alignment: human instincts as code

Can human social instincts be coded into AGI to avoid misalignment?

Deep Dive

Steven Byrnes' latest alignment post tackles the unsolved problem of designing AGI motivations from scratch. His approach: start from human innate social instincts—like approval-seeking, norm enforcement, and pride—and find ways to implement them in a brain-like AGI. He argues that if humans can steer toward good outcomes, sufficiently human-like AGIs should be able to as well, provided they have prosocial drives. The post covers three failure modes: 1) the AGI might adopt the wrong moral circle (e.g., valuing only its creator's preferences), 2) norm-following may fail without a balance of power (AGI could override human norms), and 3) consequentialist reasoning could suppress social instincts over time.

Byrnes then examines how virtues emerge: through 'person-first' identification (copying admired individuals) or 'desire-first' (deep personal passion leading to self-image changes). He suggests the desire-first pathway is crucial for truth-seeking traits in AGI. The second half of the post shifts to backward reasoning from desiderata: technical constraints (code must be writable), strategic constraints (resilience against misaligned ASIs), and societal buy-in (the plan must sound reasonable). This is an early-stage, interconnected dump of ideas, seeking feedback and deconfusion.

Key Points
  • Byrnes proposes using human social instincts (approval, pride, norm enforcement) as inspiration for AGI motivation systems.
  • Three failure modes identified: wrong moral circles, power-imbalance breaking norms, and consequentialist desires overriding social drives.
  • Desire-first pathway (deep passion → self-image) is key for fostering truth-seeking virtues in brain-like AGI.

Why It Matters

Offers a novel alignment blueprint using human-like social drives, potentially safer than pure reward-maximization.

📬 Get the top 10 AI stories daily