Robotics

Xiaomi's VLA model beats SOTA with 100K hours of real-world training

Xiaomi trains robot foundation model on 100K hours of real manipulation data.

Deep Dive

Xiaomi Robotics has unveiled Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model that can perform diverse mobile manipulation tasks in unseen environments by following natural language instructions. The model was trained on an unprecedented dataset of over 100,000 hours of real-world manipulation trajectories collected using UMI (Universal Manipulation Interface) devices. To enable rich language conditioning, the team developed a scalable auto-labeling pipeline that automatically annotates trajectory clips with natural language descriptions of scene state transitions. The training follows a two-stage recipe: pre-training on this massive dataset to imbue generalizable action-generation capabilities, followed by post-training that aligns the model with specific robot embodiments and imperative human instructions. The results demonstrate strong scaling behavior—performance consistently improves with more data and larger model sizes, and this scaling transfers directly to real-world robot deployment. On simulation benchmarks, Xiaomi-Robotics-1 sets new state-of-the-art records: a 57.6% success rate on RoboCasa365 (beating the previous 46.6%) and an average score of 20.07 on RoboDojo (compared to 13.07). The system can also be efficiently fine-tuned on complex dexterous tasks with minimal data, positioning it as a strong robot foundation policy. Code and model checkpoints will be released open-source.

The implications for robotics are significant. Xiaomi-Robotics-1 demonstrates that large-scale real-world data collection, combined with automated language annotation, can produce generalist robot models that work out-of-the-box in unseen environments. The two-stage training framework provides a blueprint for other labs: first learn broad motor skills from diverse trajectories, then fine-tune for specific embodiments and instructions. The open-source release will accelerate research in VLA models and robot foundation policies by providing both a trained model and the data pipeline. While the model currently handles mobile manipulation tasks, the approach is designed to scale to more complex behaviors. For tech professionals tracking AI and robotics, this signals that the era of general-purpose robot models trained on internet-scale real data is approaching maturity, with Chinese manufacturers like Xiaomi now leading in data scale and real-world performance. The benchmarks show a clear >23% improvement on RoboCasa and >50% on RoboDojo, indicating that real-world training data remains a bottleneck—and Xiaomi has just cranked it open.

Key Points
  • Over 100K hours of real-world manipulation trajectories collected via UMI devices
  • Achieves 57.6% success on RoboCasa365 (beating previous best of 46.6%)
  • Two-stage training: pre-training with auto-labeled language annotations, post-training for embodiment alignment

Why It Matters

Real-world data at scale yields generalist robot models outperforming prior SOTA by over 20%.

📬 Get the top 10 AI stories daily