AI Safety

LessWrong post argues Agent Foundations is either a paradigm or nothing – no middle ground

New philosophical critique says calling for 'empirical grounding' in AF is a meaningless compromise.

Deep Dive

The post examines two possible outcomes for Agent Foundations (AF): either it becomes a well-defined, self-contained network of concepts and proofs about agency (like computability theory), or it is fundamentally ill-defined – pointing at nothing, dissolving into other tasks, or only meaningful within a broader framework. The author argues that these are the only logical possibilities, ruling out the popular middle ground that 'AF needs more empirical grounding.'

If AF is coherent, empirical grounding is merely implementation, not part of the theory itself. If it is not coherent, empirical work should target whatever remains after the confusing framing is dissolved. The post uses analogies: building computers doesn't improve computability theory, and asking for empirical grounding of a non-existent field is like trying to mathematize the Christian God. For AI safety researchers, this forces a choice: either commit to developing AF as a rigorous paradigm, or abandon it for more grounded approaches.

Key Points
  • Agent Foundations (AF) aims to conceptually understand agency; either it becomes a self-contained paradigm or it is ill-defined (like 'proving God's existence mathematically').
  • The middle position 'AF needs more empirical grounding' is a 'valley of compromise' – if AF is coherent, grounding is mere implementation; if not, empirical effort is misdirected.
  • Future textbooks will either present a unified theory of agency or not; the debate cannot be resolved by adding empirical work to a conceptually confused field.

Why It Matters

For AI safety researchers, this argument demands clarity on whether Agent Foundations is a viable paradigm or a distraction.

📬 Get the top 10 AI stories daily