Robotics

AgenticFocus: Skoltech's AI turns human POV video into robot training data

New MR pipeline recovers hidden objects and hand motion from first-person video for humanoid learning.

Deep Dive

A team from Skoltech (led by Iaroslav Kolomiets) has introduced AgenticFocus, a mixed reality synthesis pipeline designed to turn standard human first-person-view (FPV) videos into training data for humanoid robots. The core problem AgenticFocus solves is that current humanoid learning pipelines struggle with hand-object occlusion, rely on oversimplified motion models, or require expensive specialized capture hardware. By restoring occluded object geometry, reconstructing full-hand motion with high fidelity, and retargeting it to a humanoid embodiment through camera-relative alignment and layered compositing, AgenticFocus produces datasets pairing focused visual observations with synchronized robot actions and states.

In benchmarks against cross-embodiment baselines, AgenticFocus demonstrated lower trajectory error and smoother wrist motion. Its SPARC scores (a metric for action quality) reached -5.18, outperforming baselines at -5.56 and -6.05. The system can leverage everyday human videos—like someone performing a dexterous task filmed from a head-mounted camera—and convert them into robot-trainable demonstrations without any special motion capture suits or multi-camera setups. This scaling property makes AgenticFocus a practical step toward generalizable humanoid skill acquisition from abundant online video data.

Key Points
  • Restores occluded object geometry from ordinary FPV human videos using mixed reality compositing.
  • Reconstructs full-hand motion and retargets it to humanoid robots with camera-relative alignment.
  • Achieves SPARC scores of -5.18, beating baseline methods at -5.56 and -6.05 for action quality.

Why It Matters

Enables scalable humanoid training from everyday human video, replacing expensive motion capture with casual recordings.

📬 Get the top 10 AI stories daily