Robots Are Now Learning New Skills by Watching Human Videos
This could make helpful home robots years cheaper — and years closer to your kitchen.
Researchers asked whether human videos can provide effective, scalable supervision for pretraining vision-language-action policies, and built a robotization pipeline that converts heterogeneous human videos into robot-aligned observations and action trajectories while inferring missing intermediate signals across annotation levels. From five human-video sources, they constructed the HuRo dataset: about 630K robotized episodes and 142M processed frames. Across four real-world manipulation tasks, increasing robotized pretraining scale improved overall completion from 51.5% to 80.3%, and OOD completion under spatial and visual shifts from 34.9% to 72.2%. Ablations showed visual robotization improves OOD robustness, and end-to-end pretraining with retargeted actions outperforms visual-only transfer. Accepted at CoRL 2026, with code and data released on the project website.
- Instead of paying for real robot practice, the team turned 142 million frames of human video into about 630,000 robot training sessions.
- Robot success on four real tasks rose from 51.5% to 80.3%, and from 34.9% to 72.2% when objects or lighting changed.
- The team released their code and data publicly, so other labs can build on it rather than starting from scratch.
Why It Matters
Cheaper robot training means useful household and warehouse robots could arrive sooner and cost far less.