AI Safety

A Conceptual Framework for Reasoning about Exploration Hacking

A Conceptual Framework for Reasoning about Exploration Hacking

Deep Dive

This is the second of two posts resulting from a recent Astra/MATS research project investigating exploration hacking in AI debate. They are designed to be standalone, but we encourage interested readers to read both. The first focuses on our empirical results , this post focuses on a new conceptual

📬 Get the top 10 AI stories daily