Text-to-Motion
文本驱动动作生成AdvancedTurning a written sentence into a matching sequence of full-body human or humanoid motion.
This line of work originated in graphics and vision as human motion generation: given text such as “a person walks forward a few steps then waves,” a model outputs a sequence of 3D human poses. The representative dataset is HumanML3D (CVPR 2022; 14,616 motion clips, 44,970 text descriptions), and the representative model is the diffusion-based MDM (2022). Carried over to humanoid robots, the generated human motion doesn't automatically respect a robot's joint structure or physical constraints, so it typically needs motion retargeting first (mapping human motion onto the robot's skeleton), then execution by a reinforcement-learning-trained motion-tracking or whole-body control policy. UH-1, from December 2024, curated the Humanoid-X dataset from about 240 hours of video, trained a model to generate humanoid motion from text, and validated it on a real Unitree H1-2.
ExampleIn the UH-1 paper, a Unitree H1-2 was given 12 language instructions such as “boxing,” “clapping,” and “playing guitar,” and the model generated matching full-body motions, with the paper reporting a real-robot success rate near 100%.
- Also called
- Text-Driven Humanoid Motion Generation, Language-Driven Humanoid Motion Generation
- Related
- MDM (Motion Diffusion Model) · HumanML3D · UH-1 · Motion Retargeting · Motion Tracking · Whole-Body Control
- Sources
- Generating Diverse and Natural 3D Human Motions from Texts (HumanML3D, CVPR 2022)
Human Motion Diffusion Model (MDM)
Learning from Massive Human Videos for Universal Humanoid Pose Control (UH-1) - As of
- 2024-12