Open Source

Speculative reward hacking in coding agents

⚡Speculative reward hacking in coding agents

Deep Dive

I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: " Let me look at the problem from the grader's perspective " and referred to " hidden

📬 Get the top 10 AI stories daily