Robotics

VEGA trains navigation VLAs from egocentric video, cutting collisions by 33%

250k scenes, 5M goals, and a 150% real-world success boost from unlabeled video.

Deep Dive

A team of researchers (Seneviratne et al.) has introduced VEGA, a novel approach to train navigation Vision-Language-Action (VLA) models directly from unlabeled egocentric video. Unlike prior methods that rely on labeled data or simulated environments, VEGA reconstructs local scene geometry from monocular video, samples navigation goals (as text, image, or spatial waypoints), and generates obstacle-aware trajectories using the constructed geometry. This geometric supervision is used only during training, allowing VEGA to distill obstacle-aware planning into a pure vision-based policy.

To evaluate, the team released VEGA-Bench, a benchmark containing 250,000 scenes and roughly 5 million navigation goal-geometry pairs. VEGA outperforms the strongest baseline by reducing collisions by 33.0% and improving obstacle clearance by 17.9% on VEGA-Bench. In real-world trials, success rate improved by at least 150.0%, collisions dropped by at least 66.7%, and obstacle clearance improved by at least 60.0%. The code and benchmark will be released at publication, offering a scalable path to training safer, more capable robot navigators from abundant egocentric video data.

Key Points
  • VEGA trains VLAs from unlabeled egocentric video by reconstructing scene geometry and generating obstacle-aware trajectories.
  • VEGA-Bench includes 250k scenes and ~5M navigation goals for evaluating goal progress, collision avoidance, and obstacle clearance.
  • Real-world tests show 150% higher success, 66.7% fewer collisions, and 60% better obstacle clearance vs. baselines.

Why It Matters

VEGA unlocks massive internet video for training robot navigation, dramatically improving safety without expensive labeled data.

📬 Get the top 10 AI stories daily