Robotics

New AI framework lets Unitree G1 robots dance to music and follow speech commands

Robots can now autonomously select movements in real-time from music or speech cues.

Deep Dive

A team led by J. M. A. Marcelo and colleagues published a paper presenting a novel multi-modal orchestration framework for dynamic humanoid whole-body control driven by semantic audio understanding. The system processes continuous audio streams and routes them into two distinct branches: music and speech. For music input, it uses audio fingerprinting and semantic embeddings to retrieve track identity and temporal alignment, dynamically mapping musical segments to pre-learned motion policies. For speech input, the system grounds commands into a discrete library of imitation-learned skills, enabling direct human-robot interaction. Both modalities share a unified interface that schedules skill execution over a reinforcement learning control pipeline.

The framework was validated in simulation and on a real Unitree G1 humanoid, demonstrating robust sim-to-real transfer and consistent audio-conditioned policy selection. This approach moves beyond pre-scripted sequences or externally triggered behaviors, providing robots with autonomy and responsiveness to dynamic audio environments. The paper was accepted at the 29th Robocup International Symposium (2026, Incheon, South Korea) and is available on arXiv. The work shows significant potential for applications in entertainment, service robotics, and collaborative human-robot environments where natural audio cues are essential.

Key Points
  • Framework processes audio into music (via fingerprinting/embeddings) or speech branches
  • Validated on Unitree G1 humanoid with robust sim-to-real transfer
  • Accepted at Robocup 2026 International Symposium

Why It Matters

Enables humanoid robots to respond dynamically to audio cues, moving beyond pre-scripted sequences for more natural interaction.

📬 Get the top 10 AI stories daily