Embodied AI Glossary中文

UH-1

UH-1 / Humanoid-XAdvanced

A large model that learns from massive internet human videos and generates humanoid robot motion from text instructions.

UH-1 was released in December 2024 by the University of Southern California (Yue Wang's group), UC Berkeley, and Toyota Research Institute, later given an oral presentation at Humanoids 2025. Humanoid robot data mostly comes from reinforcement learning and teleoperation, which is hard to scale up; this work instead learns from internet human videos. The team first built the Humanoid-X dataset: 3D human poses are extracted from videos and automatically captioned, then retargeted into humanoid keypoints and actions, yielding about 164,000 motion clips and more than 20 million robot poses. UH-1 is then trained on this: humanoid motions are first discretized into tokens, and a Transformer autoregressively generates motion tokens conditioned on a text instruction. The output can either be keypoints, which are handed to a goal-conditioned policy for tracking, or robot actions directly, executed open-loop. UH-1 demonstrates a route for scaling up humanoid data by turning human video plus text into humanoid motion.

ExampleGiven the text 'wave hello,' UH-1 generates a corresponding sequence of humanoid motion, which is handed to a lower-level controller to drive the humanoid robot's wave.

Also called
Universal Humanoid Pose Control
Related
Humanoid-X · Human Video Data · Motion Retargeting · Text-to-Motion · Action Tokenizer · Humanoid Robot
Sources
Learning from Massive Human Videos for Universal Humanoid Pose Control (arXiv 2412.14172)
UH-1 project page (PSI Lab)
As of
2024-12

See it in the full glossary →