Embodied AI Glossary中文

Humanoid-X

Humanoid-X 数据集Advanced

A humanoid-motion dataset built by extracting and captioning motion from massive amounts of internet human videos.

Humanoid-X is a dataset built by the University of Southern California, UC Berkeley, and the Toyota Research Institute, described in the paper Learning from Massive Human Videos for Universal Humanoid Pose Control, made public in December 2024, with the paper as an oral presentation at Humanoids 2025. The pipeline: mine human-motion videos from the internet, automatically generate text descriptions, estimate the 3D pose of the person in each video, retarget that motion into joint targets for a humanoid robot, and train a control policy to turn those targets into motion the robot can actually execute. The final dataset has 163,800 samples and over 20 million humanoid robot poses, each entry carrying the source video, text, human pose, robot keypoints, and the resulting action. The UH-1 model trained on it takes a text instruction as input and outputs humanoid robot motion. Its significance is bypassing expensive teleoperation and motion capture altogether, expanding humanoid motion data directly from human video already available online.

ExampleGiven the instruction “wave hello” as input, UH-1 outputs a sequence of humanoid robot joint motions that execute a wave, in simulation or on a real robot.

Also called
UH-1 Dataset, Learning from Massive Human Videos for Universal Humanoid Pose Control
Related
Human Video Data · Internet Video Data · Motion Retargeting · Text-to-Motion · Humanoid Robot · UH-1
Sources
Learning from Massive Human Videos for Universal Humanoid Pose Control (arXiv)
UH-1 项目主页(PSI Lab) (Chinese)
As of
2025-10

See it in the full glossary →