Embodied AI Glossary中文

MiMo-Embodied (Xiaomi)

小米 MiMo-EmbodiedAdvanced

An open-source 7B vision-language model from Xiaomi covering both autonomous-driving and embodied-AI understanding and planning.

MiMo-Embodied is a vision-language model (VLM) whose technical report was released, and weights open-sourced, by Xiaomi's embodied-AI team in November 2025; it has about 7 billion parameters, with weights public on Hugging Face. The premise is that autonomous driving and indoor robots both need spatial understanding and planning, yet are usually trained entirely separately. MiMo-Embodied puts both kinds of data into a single model, trained through multiple stages with curated data plus chain-of-thought and reinforcement-learning fine-tuning, covering affordance prediction, task planning, and spatial understanding on the embodied side, and environment perception, state prediction, and driving planning on the driving side. The report says it matches or beats comparable open- and closed-source models across 17 embodied-AI benchmarks and 12 autonomous-driving benchmarks, and observes positive transfer between the two domains, each reinforcing the other. It evaluates understanding, reasoning, and planning ability; it does not directly output a robot's joint actions itself.

ExampleShow it a kitchen photo and ask “where should the cup be grasped,” and it can point out the graspable location on the image; show it driving footage, and it can describe the state of surrounding vehicles and give a next driving plan.

Also called
MiMo-Embodied-7B, X-Embodied Foundation Model
Related
Vision-Language Model · Embodied Reasoning · Affordance · Spatial Reasoning · Autonomous Driving · Cross-Embodiment
Sources
MiMo-Embodied: X-Embodied Foundation Model Technical Report (arXiv 2511.16518)
XiaomiMiMo/MiMo-Embodied (GitHub)
As of
2026-04

See it in the full glossary →