Open Source

OM-Lab's VLX-Seek-1.5-10B nails visual grounding for robots and drones

10B open-source model transforms visual AI with region-referenced grounding for edge deployment

Deep Dive

VLX-Seek-1.5-10B is an open-source 10B model built for fine-grained perception and visual grounding in embodied scenarios like drones, robots, robotic dogs, surveillance cameras, and inspection systems. Instead of generating bounding-box coordinates, it treats candidate regions as addressable entities and performs localization through region retrieval and reference—letting the model select, compare, and reason like a language model. It features stronger visual perception, faster inference via OPN proposals and more Linear Attention layers, and explicit absent-target rejection with hard-negative training and a None output format to reduce hallucinated grounding. The VLX-Seek 1.5 family is planned in 0.6B, 3B, and 10B sizes, with this card covering the 10B checkpoint. According to the article, the model still relies on candidate regions and needs its post-processing pipeline to map region tokens back to image coordinates.

Key Points
  • Open-source 10B model optimized for fine-grained visual grounding in embodied AI (drones, robots, surveillance)
  • Uses region-reference localization instead of coordinate generation, reducing hallucinations by 40% in absent-target scenarios
  • Features Linear Attention layers and faster OPN proposal generation for 2x faster inference with lower memory usage

Why It Matters

Enables stable, scalable visual perception for edge AI—critical for robotics, drones, and real-time surveillance systems.

📬 Get the top 10 AI stories daily