AI Safety

Anthropic probability fails under duplication events

Duplicating an AI agent breaks probability theory. Here's why.

Deep Dive

Stuart Armstrong, a researcher at LessWrong, has published an impossibility result showing that standard probability theory collapses when agents are duplicated. The Duplicate Sleeping Beauty thought experiment illustrates the problem: if an initial agent is either awakened alone or duplicated into two identical copies based on a coin flip, no probability function can remain consistent across both outcomes.

The core issue lies in the martingale condition (conservation of expected evidence) and simple Bayes updating. Both principles fail in duplication scenarios, forcing a choice between violating Bayesian consistency or the martingale property. Armstrong argues SIA (Self-Indication Assumption) is the most resilient framework, though it isn't perfect either. He previews D-SIA (Distributional SIA) as a potential solution to these infinity-related problems in his upcoming final post.

Key Points
  • Anthropic probability fails when existing agents are duplicated (not just created), per Stuart Armstrong's new research
  • Duplicate Sleeping Beauty thought experiment shows no probability function can satisfy both martingale and Bayes consistency simultaneously
  • SIA (Self-Indication Assumption) emerges as the most robust candidate, but Armstrong will propose D-SIA to address its infinity issues

Why It Matters

This work fundamentally challenges AI safety assumptions about agent duplication and probability consistency in distributed systems.

📬 Get the top 10 AI stories daily