Alibaba's Qwen-Robot Series: Three Models That Unite Language and Action — And One Can Predict What Happens Next
Over 38,000 hours of training data for robot manipulation across multiple platforms.
Alibaba's Qwen team launched the Qwen-Robot series, a robotics suite of three foundation models that bridge language understanding with physical actions. Qwen-RobotNav extends vision-language capabilities into mobile robotics with controllable observation encoding and tool-based interfaces, unifying instruction following, goal-directed navigation, target tracking, and autonomous driving within a single framework. This allows robots to interpret natural language commands and navigate complex environments.
Qwen-RobotManip standardizes state-action space by representing end-effector motion as incremental poses in the camera coordinate system. Trained on over 38,100 hours of fully open-source data, it supports large-scale learning across multiple robot platforms for diverse manipulation tasks. Qwen-RobotWorld serves as a general-purpose world model, connecting vision-language understanding with future-state prediction through a natural-language action interface, forecasting physically consistent outcomes in navigation, driving, and manipulation—enabling generalization across diverse embodied AI tasks.
- Qwen-RobotNav unifies instruction following, navigation, tracking, and autonomous driving in one mobile robotics model.
- Qwen-RobotManip is trained on 38,100+ hours of open-source data for cross-platform robot manipulation.
- Qwen-RobotWorld predicts future physical states using a natural-language action interface, generalizing across navigation, driving, and manipulation scenarios.
Why It Matters
Alibaba's models advance embodied AI by aligning language with robotics, enabling versatile, real-world automation.