VLMs help robots adapt to dynamic group formations with 25% fewer collisions
Robots can now naturally accompany groups that change formations, thanks to AI reasoning.
Robots that accompany groups of humans have long struggled with dynamic formations—people don't walk in fixed patterns. A new paper from Cong-Thanh Vu and Yen-Chen Liu at National Taiwan University tackles this with an Adaptive Companionship method powered by Vision-Language Models (VLMs). The approach first detects group members, then uses a perceptual module to generate visual representations of the interaction space as input to the VLM, which infers appropriate companion positions and understands group dynamics. This semantic reasoning is combined with a Model Predictive Path Integral (MPPI) controller to ensure smooth, collision-free navigation.
Evaluated across five varied scenarios—including walking, stopping, splitting, and re-grouping—the method achieved a 15% higher success rate and a 25% reduction in collisions compared to baseline methods. A user study confirmed that the robot's companionship behaviors felt natural and socially appropriate. The work, accepted to IEEE/RSJ IROS 2026, demonstrates that coupling large language model reasoning with classical control can bridge the gap between rigid robotic following and fluid human social behavior. Real-world applications include guide robots, hospital assistants, and personal companions.
- Combines Vision-Language Models (VLMs) with a Model Predictive Path Integral (MPPI) controller for safe navigation.
- Demonstrated 15% improvement in success rate and 25% reduction in collision rate across 5 dynamic group scenarios.
- User study rated robot companionship behaviors as natural and socially appropriate; accepted to IROS 2026.
Why It Matters
Enables social robots to naturally follow human groups in real-world settings, cutting collisions by 25%.