Embodied AI Glossary中文

Qwen-Robot Series

千问 Qwen-Robot 系列Common

Alibaba's Qwen team released this trio of 2026 embodied models, one each for manipulation, navigation, and world modeling.

Qwen-Robot is a group of embodied models Alibaba's Qwen team released together in June 2026, all built on the Qwen vision-language model, with each of the three members handling a different job. Qwen-RobotManip is a manipulation VLA: a Qwen-VL backbone followed by a flow-matching DiT action head, trained only on open-source robot data plus robot trajectories synthesized from human-hand videos, totaling about 38,100 hours of pretraining data, with an emphasis on aligning data across different robot embodiments before scaling up. Qwen-RobotNav uses a unified waypoint-prediction interface to handle vision-language navigation, object search, target tracking, and autonomous driving all at once. Qwen-RobotWorld is a language-conditioned video world model used to synthesize training data and evaluate policies. The official repository states there is currently no plan to release weights for Manip or Nav.

ExampleQwen-RobotNav was deployed zero-shot on a Unitree Go2 quadruped, running inference on a Jetson Thor at about 5Hz to navigate unfamiliar environments by language instruction.

Also called
Qwen-RobotManip, Qwen-RobotNav, Qwen-RobotWorld
Related
Qwen-VL · Vision-Language-Action Model · World Model · Vision-and-Language Navigation · Flow Matching · Cross-Embodiment
Sources
Qwen-RobotManip Technical Report (arXiv:2606.17846)
QwenLM/Qwen-RobotNav GitHub 仓库 (Chinese)
Qwen-RobotWorld Technical Report (arXiv:2606.17030)
As of
2026-09

See it in the full glossary →