New AI Benchmark Teaches Robots Who Does What to Whom
Future home robots could finally tell who's helping whom — not just what objects are in the room.
Most AI vision systems describe a scene like a stack of separate sticky notes: person, cup, holding. Useful, but incomplete. They miss the structure of an event — who started the action, who received it, and whether the same person played both parts. HEIR (Human-Entity Interactions with Functional Roles) is a new public benchmark designed to fix that. It contains 18,730 images labeled with six functional roles, 105 actions, and 437 nouns, covering interactions with objects, between people, and people acting on themselves.
The real world is messy, and the dataset proves it: 51.6% of the images contain more than one person acting, and 62.1% contain more than one action happening at once. The same face can appear in several events, and someone can switch from helper to receiver between photos. Old scoring methods graded each link one at a time, so a photo where a person hands a cup to another person could score well even if the model got the direction backwards — like grading a soccer match by counting passes without ever checking who scored.
The team also built a new AI method called CoRISP that ties the same people together across images and predicts a complete, normalized set of participants and roles for each event. It uses two clever signals: how many people are involved, and how many people can fill the same role. Across 16 competing systems, CoRISP came out ahead on the hardest cases — repeated roles and shared participants — by 2.87 and 3.82 points. It also hit 73.72 and 76.23 role accuracy on the well-known V-COCO test set.
Why should you care? Role-aware vision is the missing piece for robots that work alongside people — a warehouse robot needs to know who is handing it a package versus who is walking past. It could also improve video search ('find clips where someone is teaching someone else'), better descriptions for blind users, and safety cameras that distinguish play from a fight. The honest limitation: this is a research release, and widely deployed role-tracking cameras raise real privacy questions about how closely machines watch human behavior.
- HEIR is a free, public set of 18,730 images that labels not just what's happening but who is doing what to whom.
- Most images are crowded: over half have multiple people acting and over 60% have multiple actions at once.
- The team's CoRISP method beat 16 rival systems on the hardest cases, where roles repeat or people share the spotlight.
Why It Matters
Better role-reading AI means smarter home robots, easier video search, and cameras that grasp context — though also sharper privacy trade-offs.