Embodied AI Glossary中文

Xiaomi-Robotics-0

小米 Xiaomi-Robotics-0Advanced

Xiaomi's open-source, 4.7-billion-parameter VLA, built to execute actions in real time and smoothly on a consumer GPU.

Xiaomi-Robotics-0 is a vision-language-action model released and open-sourced by Xiaomi in February 2026, with 4.7 billion parameters in total, using Qwen3-VL-4B as its vision-language backbone and attaching a diffusion Transformer action expert that generates actions via flow matching. Pretraining used roughly 200 million timesteps of cross-embodiment robot trajectories (from DROID, MolmoAct data, and Xiaomi's own collected data), plus more than 80 million vision-language samples. Its main focus is fixing the action stutter caused by slow inference: it uses asynchronous execution, inferring the next action segment while still executing the previous one, and feeds already-committed actions back into the model as a prefix, paired with a Λ-shaped attention mask to keep the transition smooth; on an RTX 4090, inference latency is about 80 milliseconds. In July 2026, Xiaomi followed up with Xiaomi-Robotics-1, trained on more than 100,000 hours of real-robot data.

ExampleThe company reports 98.7% average success on LIBERO and an average of 4.75 consecutively completed tasks on CALVIN ABC→D; on a real robot, it demonstrated two bimanual tasks — disassembling LEGO and folding a towel.

Also called
Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
Related
Vision-Language-Action Model · Asynchronous Inference · Real-Time Chunking · Action Expert · Xiaomi · MiMo-Embodied (Xiaomi)
Sources
Xiaomi-Robotics-0 (arXiv:2602.12684)
Xiaomi-Robotics-0 技术报告 HTML 版 (Chinese)
Xiaomi-Robotics-1 (arXiv:2607.15330)
As of
2026-07

See it in the full glossary →