Embodied AI Glossary中文

Text-to-Motion

文本驱动动作生成Advanced

Turning a written sentence into a matching sequence of full-body human or humanoid motion.

This line of work originated in graphics and vision as human motion generation: given text such as “a person walks forward a few steps then waves,” a model outputs a sequence of 3D human poses. The representative dataset is HumanML3D (CVPR 2022; 14,616 motion clips, 44,970 text descriptions), and the representative model is the diffusion-based MDM (2022). Carried over to humanoid robots, the generated human motion doesn't automatically respect a robot's joint structure or physical constraints, so it typically needs motion retargeting first (mapping human motion onto the robot's skeleton), then execution by a reinforcement-learning-trained motion-tracking or whole-body control policy. UH-1, from December 2024, curated the Humanoid-X dataset from about 240 hours of video, trained a model to generate humanoid motion from text, and validated it on a real Unitree H1-2.

ExampleIn the UH-1 paper, a Unitree H1-2 was given 12 language instructions such as “boxing,” “clapping,” and “playing guitar,” and the model generated matching full-body motions, with the paper reporting a real-robot success rate near 100%.

Also called
Text-Driven Humanoid Motion Generation, Language-Driven Humanoid Motion Generation
Related
MDM (Motion Diffusion Model) · HumanML3D · UH-1 · Motion Retargeting · Motion Tracking · Whole-Body Control
Sources
Generating Diverse and Natural 3D Human Motions from Texts (HumanML3D, CVPR 2022)
Human Motion Diffusion Model (MDM)
Learning from Massive Human Videos for Universal Humanoid Pose Control (UH-1)
As of
2024-12

See it in the full glossary →