Robotics

HRIBench benchmark exposes VLA robots' collaboration gap, boosts real-world success 4x

New benchmark reveals even top VLA models fail at basic human-robot coordination tasks.

Deep Dive

HRIBench, developed by a team at Renmin University of China, tackles a critical blind spot in robotics: most vision-language-action (VLA) benchmarks measure isolated manipulation skills, ignoring the messy reality of human-robot collaboration. The new benchmark explicitly models shared agency through structured scenario scripts that define agent roles, temporal dependencies, coordination constraints, and human behavior distributions. It introduces three interaction archetypes – Instructor, Collaborator, and Intruder – to test intent communication, joint coordination, and robustness under human intervention. With 13 tasks and over 650 evaluation episodes generated from diverse trajectories and scene variations, HRIBench provides interpretable metrics beyond binary success: synchronization, responsiveness, protocol compliance, and safety.

When the team evaluated adapted policies based on GR00T, pi0.5, and ACT under a unified protocol, the results were striking: current foundation robot policies struggle substantially in collaborative settings despite strong manipulation ability. Fine-tuning on the benchmark consistently improved collaborative performance. In a real-world adaptation study, simulation data from HRIBench boosted GR00T N1.5's physical-task success rate from a mere 0.10 to 0.43 – a 4.3x improvement. This demonstrates the benchmark's practical value for advancing interaction-centric robot learning and highlights a critical direction: robots must learn to coordinate, not just manipulate.

Key Points
  • HRIBench introduces 13 role-conditioned tasks with 3 interaction roles (Instructor, Collaborator, Intruder) and over 650 evaluation episodes.
  • Current VLA models (GR00T, pi0.5, ACT) show major failures in temporal coordination and intent-aware behavior despite strong manipulation skills.
  • Fine-tuning on HRIBench boosted real-world task success for GR00T N1.5 from 0.10 to 0.43, a 4.3x improvement.

Why It Matters

This benchmark fills a crucial gap in robotics evaluation, pushing research toward truly collaborative human-robot systems.

📬 Get the top 10 AI stories daily