Robotics

VoLoAgent uses VLM to orchestrate robots for open-vocabulary long-horizon tasks

New system handles flexible instructions and complex scenes with interruptible tools.

Deep Dive

VoLoAgent, developed by a team including Siyi Chen, Hugo Hadfield, and researchers from NVIDIA and academic institutions, tackles the challenge of open-vocabulary long-horizon manipulation in robotics. Traditional systems struggle with flexible instructions and complex multi-object scenes because they cannot adapt mid-task or recover from failures. VoLoAgent addresses this by implementing a closed agent loop where a vision-language model (VLM) acts as a physical orchestrator. It treats a vision-language-action model (VLA) or whole-arm manipulation (WAM) as an interruptible tool that can be steered in real-time, alongside other vision models and action primitives. This design accounts for the critical timing constraints of physical environments, where reasoning cannot pause execution.

To evaluate these capabilities, the team introduces RoboVoLo, a high-fidelity benchmark that tests open-vocabulary long-horizon manipulation across categories like common sense, memory/state tracking, complex references, and world knowledge. The benchmark includes both task-level success metrics and failure-mode diagnostics. Experimental results show VoLoAgent substantially outperforms single VLA/VLM systems or naive tool-based approaches. Real-robot experiments validate these findings, demonstrating that the system can plan, execute, monitor, and recover from failures autonomously. This work represents a significant step toward more capable and flexible robotic manipulation in unstructured environments.

Key Points
  • VoLoAgent uses a VLM to orchestrate heterogeneous robot capabilities (VLA/WAM, vision models, action primitives) as interruptible tools for long-horizon tasks.
  • Introduces RoboVoLo benchmark with task-level success and failure diagnostics across common sense, memory, complex references, and world knowledge.
  • Substantially outperforms single VLA/VLM or tool-based systems, validated on real-robot experiments.

Why It Matters

Enables robots to understand flexible instructions and recover from failures autonomously in complex environments.

📬 Get the top 10 AI stories daily