Research & Papers

DROPJ: New method trains safe AI agents using human preferences and world models

Human justifications boost AI safety without complex reward engineering.

Deep Dive

Researchers from the University of Southampton and King's College London have introduced DROPJ (Deployment via Reward Optimization with Preferences and Justifications), a novel human-centered framework for training safe AI agents in environments with unknown dynamics and no predefined reward function. The method first learns a world model—a learned simulator—from a dataset of prior real-world trajectories. A human then interacts with this simulator to generate informative simulated trajectories. From these, the system elicits pairwise preferences over trajectory segments and crucially, also collects a textual justification for each choice. These justified preferences are used to train a reward model, which is then combined with the world model to directly deploy the agent using model predictive control (MPC).

In real-user experiments, the team found that generating informative simulated trajectories significantly reduces computational cost during training compared to other human-feedback strategies. Moreover, using preferences over other types of feedback (like scalar ratings) substantially improves deployment performance. Most importantly, incorporating safety justifications alongside preferences allows the agent to prioritize user-prescribed aspects of safety during deployment—enhancing overall safety without sacrificing performance. The paper, presented at ICAART 2026, demonstrates that DROPJ offers a practical path to aligning AI behavior with human safety values without needing complex reward engineering or full environment models.

Key Points
  • DROPJ uses a world model learned from real-world data to reduce unsafe exploration during training.
  • Human preferences paired with textual justifications improve both performance and safety in deployment.
  • Informative simulated trajectories cut computational cost compared to other human feedback methods.

Why It Matters

Makes safe AI deployment practical by replacing hard-to-design reward functions with intuitive human input.

📬 Get the top 10 AI stories daily