AI Safety

LessWrong post proposes extended science to solve alignment's ought problem

⚑PeterL suggests adding 'preference' to fundamental physics to derive morality.

Deep Dive

PeterL's LessWrong post revisits the classic philosophical 'ought' problem as a root cause of AI alignment difficulties. Standard science describes what is (facts), but cannot prescribe what ought to be (values). PeterL argues this gap stems from the current scientific method's exclusive reliance on necessity. He proposes an 'extended science' where fundamental building blocks include both necessity and preference. The description of these building blocks would be optimized not only for simplicity and fit to data, but also for a preference-to-necessity complexity ratioβ€”i.e., what can be explained by preference should not be explained by necessity.

This extended science would provide a proto-criterion for evaluating 'better' states (where better = more preferred). From there, morality could be framed as a justice principle across all fundamental building blocks applied to that criterion. PeterL admits this is just an idea and keeps the post short. While highly speculative, the proposal resonates with ongoing debates in AI alignment about how to ground ethics in substrate-independent principles, potentially offering a fresh angle for value learning or coherent extrapolated volition approaches. No implementation details or benchmarks are provided.

Key Points
  • Identifies the 'ought' problem as central to AI alignment: facts cannot imply values.
  • Proposes 'extended science' adding a preference category to fundamental physics.
  • Uses preference-to-necessity complexity ratio to select better states as a moral proto-criterion.

Why It Matters

A novel philosophical framework that could influence how researchers approach grounding AI values in fundamental physics.

πŸ“¬ Get the top 10 AI stories daily