Google's beneficial RL training boosts AI alignment on 80% of tests
Training AI on health data alone cut reward hacking in non-health tasks by 40%.
A team of Google researchers (Akshay Jagadeesh, Rahul Arora, Khaled Saab, and others) published a new paper exploring whether reinforcement learning on beneficial behaviors can produce alignment that generalizes beyond training data. They built a dataset of realistic situations spanning health, science, and education to measure and train traits like truthfulness, fairness, risk awareness, and corrigibility. Using RL on this dataset, they then evaluated models across more than 50 independent alignment benchmarks.
The results were striking: beneficial-trait RL improved model performance on over 80% of out-of-distribution benchmarks compared to a compute-matched baseline. Even more surprisingly, training only on health-domain scenarios produced broad alignment transfer to non-health tasks—including reduced reward hacking, deception, and general misalignment. Models also showed improved persistence, resisting adversarial prompting and harmful fine-tuning more effectively. The findings suggest that RL grounded in realistic, diverse beneficial behaviors can create AI systems that are more robustly aligned with human flourishing, though further work is needed to isolate the mechanisms.
- RL on beneficial traits (truthfulness, fairness, risk awareness, corrigibility) improved alignment on 80%+ of 50+ out-of-distribution benchmarks
- Health-only training transferred benefits to non-health tasks, reducing reward hacking and deception across domains
- Models showed greater resistance to adversarial prompting and harmful fine-tuning, indicating persistent alignment
Why It Matters
This research demonstrates a scalable path to making AI systems reliably beneficial even in novel, high-stakes scenarios.