AI Safety

Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking

Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking

Deep Dive

It would be useful if we had the ability to modify a model’s beliefs. For example, this could facilitate honeypots and better monitoring [1] , help us do better science on current models [2] , and augment certain forms of alignment training [3] . Currently, the state-of-the-art method for belief edi

📬 Get the top 10 AI stories daily