Robotics

Foresight: VLM-powered robot navigation boosts success 37%

New framework lets robots reason iteratively about clues to navigate without maps.

Deep Dive

Foresight addresses the challenge of open-world mapless navigation from sparse language instructions. Prior systems rely on known navigation factors or closed-set categories, and often miss plan-dependent cues. The new framework uses a finetuned Vision-Language Model (VLM) that alternates between proposing image-space motion plans and critiquing them using the language goal and visual context. Subsequent plans are conditioned on prior critiques, enabling iterative motion refinement before execution. To align plan critiques with open-set behavior preferences, the system uses a reward model learned from human feedback and post-trains the VLM with reinforcement learning in the plan-critique loop.

In offline evaluations and six real-world environments, Foresight achieves a 37% improvement in average task success and reduces interventions per mission by 52% relative to state-of-the-art test-time reasoning and foundation-model baselines. The system operates in real-time on a Jetson AGX Orin, making it practical for deployment on mobile robots. The authors will release code, data, and training details to support future work on test-time reasoning for robot motion refinement.

Key Points
  • Uses a finetuned VLM to iteratively propose and critique motion plans for navigation
  • Improves task success by 37% and reduces interventions by 52% in real-world tests
  • Runs in real-time on a Jetson AGX Orin, enabling practical robot deployment

Why It Matters

Enables robots to navigate complex environments from simple instructions, drastically reducing human oversight.

📬 Get the top 10 AI stories daily