AI Safety

MIRI's Garrabrant proves Dutch-book-resistant probabilities must be 'centering-uniform'

A new math framework shows Sleeping Beauty-style puzzles have a unique, non-exploitable probability assignment.

Deep Dive

A new paper formalizes consistent probability assignments over “centered worlds” — uncentered worlds tagged with a “here and now” — and shows that the probability policies resistant to Dutch book arguments are precisely those that are “centering-uniform.” This means starting with a fixed measure over uncentered worlds, weighting by the number of occurrences of the agent’s information state, and spreading probability uniformly across centers sharing the same world and information state. The framework generalizes standard Bayesian updating by assuming a causal decision theorist in the background, and it connects to anthropic thought experiments like Sleeping Beauty and duplication. Memory loss is argued to be decision-theoretically similar to being one member of a collective with shared goals: both motivate taking actions based on centered-world probabilities in a consistent way. The resulting centering-uniform assignments closely resemble Bostrom’s strong self-sampling assumption and its SIA variant, differing only by a population scaling factor. The article also shows that resistance to weak Dutch books forces the underlying measure to be non-dogmatic, assigning nonzero weight to all populated uncentered worlds — a stricter form of centering-uniformity that aligns even more extensively with SSSA/SIA.

Key Points
  • Centering-uniformity: probabilities must weight uncentered worlds by occurrence count of the agent's information state, with uniform distribution over identical centers.
  • Generalizes the Dutch book argument for Bayesian updating to centered worlds under CDT, avoiding issues with EDT bet acceptance.
  • Strict centering-uniformity (from weak Dutch book resistance) enforces close agreement with Bostrom's SSSA/SIA, relevant for Sleeping Beauty and AI memory-loss/collective scenarios.

Why It Matters

Gives AI safety and anthropic reasoning a mathematically rigorous, exploitable-proof probability framework for self-location, memory loss, and multi-agent coordination.

📬 Get the top 10 AI stories daily