Research & Papers

RENEW uses human feedback to fix AI world model exploitation in RL

Human intuition can spot when AI models hallucinate physics – RENEW leverages that.

Deep Dive

World models in offline reinforcement learning often suffer from model exploitation—producing unrealistic dynamics in regions with sparse data. Prior fixes either require expensive expert demonstrations or use conservative algorithms that limit generalization. Now, researchers from Stanford introduce RENEW, a novel framework that repairs exploitation by directly incorporating human preferences over imagined rollouts. The core idea: humans have strong intuitive physics and can easily spot egregious dynamics hallucinations.

RENEW formalizes this as Dynamics Learning from Human Feedback (DLHF), a Bradley-Terby preference loss that compares trajectory log-likelihoods under the learned dynamics model. Naive DLHF is sample-inefficient, so RENEW uses epistemic uncertainty to focus human feedback where the model is most exploitable. Experiments on Jumanji and classic control benchmarks show RENEW makes the approach practical—boosting sample efficiency, limiting catastrophic forgetting, and reducing exploitation in pretrained world models. This opens a new avenue for addressing model exploitation without costly demonstrations.

Key Points
  • RENEW uses human preferences (Bradley-Terry loss) over imagined rollouts to directly repair world model exploitation.
  • Epistemic uncertainty sampling focuses human feedback on the most exploitable regions, improving sample efficiency.
  • Tested on Jumanji and classic control tasks—reduces catastrophic forgetting and outperforms naive DLHF.

Why It Matters

Human feedback can now cheaply fix AI hallucinations in world models, enabling safer offline RL.

📬 Get the top 10 AI stories daily