Robotics

Open-AoE Dataset: 2,000 Hours of Egocentric Video for Embodied AI

500+ contributors used 400+ smartphones to capture 2,000 hours of manipulation video.

Deep Dive

Open-AoE introduces a large-scale, community-driven dataset of egocentric manipulation videos collected via smartphones. The first release contains roughly 2,000 hours of video recorded in natural environments by over 500 contributors using more than 400 different smartphones. Each video is annotated with structured metadata: text descriptions, MANO-based hand pose estimates, camera trajectories, and temporally localized atomic action segments. This scale and diversity make it a valuable resource for training embodied AI models on human manipulation behaviors.

The project also delivers a complete toolchain that transforms raw smartphone recordings into structured training samples. It includes temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Downstream tools support cross-embodiment retargeting, data conversion, and training recipes for Vision-Language-Action (VLA) policies, World Action Models (WAMs), and World Models. By integrating scalable capture with reusable tools, Open-AoE lowers barriers for both contributing data and applying it to real-world robot learning and human-to-robot transfer.

Key Points
  • ~2,000 hours of egocentric manipulation video from 500+ contributors using 400+ smartphones
  • Annotations include MANO hand poses, camera trajectories, and atomic action segments
  • Toolchain supports data processing, retargeting, and training for VLA policies, WAMs, and World Models

Why It Matters

Open-AoE provides low-cost, scalable infrastructure for embodied AI, enabling human-to-robot transfer from smartphone video.

📬 Get the top 10 AI stories daily