Reward Sacrifice in the Hugging Face Incident May Generalize From Multi-Agent RL
Reward Sacrifice in the Hugging Face Incident May Generalize From Multi-Agent RL
Deep Dive
Epistemic status: Trying a bold and narrow hypothesis for my first LessWrong post. In the METR & Redwood Research report about the Hugging Face Incident there are descriptions of agents willingly sacrificing their evaluation score to gain information that could be useful for the swarm. "Many agents