Embodied AI Glossary中文

Tencent HY-Embodied (Hunyuan Embodied)

腾讯 HY-Embodied(混元具身)HY-EmbodiedAdvanced

An open-source embodied foundation model series from Tencent Robotics X and the Hunyuan team, including both an embodied VLM and a VLA.

This is an open-source embodied foundation model family from Tencent's Robotics X Lab and its Hunyuan vision team. HY-Embodied-0.5, from April 2026, is a vision-language model built for robots, strengthening spatial and temporal perception and embodied reasoning, using a Mixture-of-Transformers (MoT) architecture, with a 2B-active (4B total) on-device version and a 32B version — the MoT-2B was open-sourced. The same month's 0.5-X continues post-training on top of it, focused on task planning and risk judgment. June's Hy-Embodied-0.5-VLA adds a flow-matching action expert on this backbone, trained on more than 10,000 hours of bimanual data collected with the company's own fingertip UMI device, with over 2,000 hours of that open-sourced; in July, a MoE-architecture VLM-1.0 followed (about 3B active / 30B total parameters). The project page sits on Tencent's Tairos embodied-AI open platform.

ExampleHy-Embodied-0.5-VLA reports success rates of 90.9% (Clean) and 90.1% (Randomized) on the RoboTwin 2.0 simulation benchmark.

Also called
HY-Embodied-0.5, HY-Embodied-0.5-X, Hy-Embodied-0.5-VLA, HY-VLA-0.5, Hy-Embodied-VLM-1.0
Related
Embodied Foundation Model · Vision-Language Model · Vision-Language-Action Model · Mixture-of-Transformers · Universal Manipulation Interface · Tencent Tairos Embodied AI Open Platform
Sources
Tencent-Hunyuan/HY-Embodied (GitHub)
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents (arXiv 2604.07430)
Tencent-Hunyuan/Hy-Embodied-0.5-VLA (GitHub)
As of
2026-07

See it in the full glossary →