Robotics

New EmbodimentSemantic dataset benchmarks VLM spatial grounding for robot manipulation

60K manipulation frames reveal VLMs struggle with depth-aware spatial relations

Deep Dive

Spatial grounding remains a critical weakness for vision-language-action (VLA) systems in robotics. While current models can recognize objects and follow instructions, they lack explicit representation of spatial arrangements like support, containment, occlusion, and depth ordering. To address this, researchers from multiple institutions introduce EmbodimentSemantic, a dataset and benchmark that represents scenes as directed object-relation-object triplets with a fixed set of spatial relations. The dataset includes real-world manipulation trajectories collected using the low-cost SO101 robot arm, providing a practical testbed for evaluating relational grounding.

To enable controlled validation, the team also built a LIBERO simulator benchmark with over 60,000 manipulation frames and more than 120,000 camera-specific scene graphs from paired third-person and wrist views. Ground-truth relations are automatically derived from MuJoCo geometry, world coordinates, camera projections, and visibility constraints. Experiments across multiple open-source and commercial vision-language models (VLMs) reveal a consistent pattern: models often predict plausible relations but struggle with exact depth-aware and viewpoint-dependent spatial structure. The work provides a unified framework for diagnosing spatial grounding in VLM perception and testing its utility for downstream robotic manipulation tasks.

Key Points
  • Dataset uses directed object-relation-object triplets for explicit spatial relation representation, enabling direct evaluation of object binding and relation prediction
  • Includes 60K manipulation frames from LIBERO simulator with 120K camera-specific scene graphs, plus real-world data from low-cost SO101 robot arm
  • Tests on VLMs reveal failure in depth-aware and viewpoint-dependent spatial structure, highlighting a key bottleneck for embodied AI

Why It Matters

Improving spatial grounding in VLMs is critical for reliable robot manipulation in real-world environments

📬 Get the top 10 AI stories daily