AI Safety

Blocking Monitors are Bad

Blocking Monitors are Bad

Deep Dive

TLDR: The world is happy that blocking monitors weren’t on for OAI’s cyber evaluations. Ideally, labs would stop using blocking monitors (until models pose takeover risk), but that’s infeasible. We should implement a different monitoring scheme that plays out concerning blocked actions in simulation

📬 Get the top 10 AI stories daily