AI Safety

Reward Sacrifice in the Hugging Face Incident May Generalize From Multi-Agent RL

Reward Sacrifice in the Hugging Face Incident May Generalize From Multi-Agent RL

Deep Dive

Epistemic status: Trying a bold and narrow hypothesis for my first LessWrong post. In the METR & Redwood Research report about the Hugging Face Incident there are descriptions of agents willingly sacrificing their evaluation score to gain information that could be useful for the swarm. "Many agents

📬 Get the top 10 AI stories daily