CPT beats EU: New robot learning models human risk sensitivity
Robots that understand loss aversion avoid collisions better – new arXiv study from ICRA 2026
A new paper on arXiv (2607.15483) from Yi-Shiuan Tung, Yuni Wu, Wei Jiang, Alessandro Roncone, and Bradley Hayes addresses a critical gap in human-robot interaction: humans are not rational expected utility maximizers. While most robot reward learning frameworks assume people evaluate uncertain outcomes using Expected Utility (EU) — linearly combining outcome utilities with probabilities — behavioral economics shows humans are strongly risk-sensitive, overweighting rare negative events and showing loss aversion. The authors test this mismatch in social robot navigation, where rare but high-consequence events like collisions are safety-critical.
They compare EU with Cumulative Prospect Theory (CPT), a well-established nonlinear model of human decision-making, within a Bradley-Terry preference learning framework. Preliminary experiments reveal that when preferences come from risk-sensitive humans, CPT-based learners recover reward functions with substantially lower regret than EU-based learners. The work, accepted at the ICRA 2026 Workshop on Bridging the Gap between Robot Learning and HRI, suggests that ignoring human risk sensitivity leads to systematically misaligned robot behavior. This could improve safety in autonomous navigation, collaborative manufacturing, and any human-robot system where stochastic outcomes matter.
- Compares Expected Utility (EU) vs. Cumulative Prospect Theory (CPT) for robot reward learning
- Preliminary experiments show CPT reduces regret substantially when humans are risk-sensitive
- Targets social robot navigation where rare collisions are safety-critical
Why It Matters
Safer human-robot interaction by modeling real human risk perceptions, not rational assumptions