HumanML3D
HumanML3D 数据集AdvancedA text-to-motion dataset pairing 14,000 clips of 3D human motion with 45,000 natural-language descriptions.
HumanML3D comes from Guo et al.'s CVPR 2022 paper, Generating Diverse and Natural 3D Human Motions From Text. The authors drew 14,616 motion clips from two human motion-capture datasets, AMASS and HumanAct12, and had people write 3–4 English sentences describing each one, yielding 44,970 sentences totaling about 28.59 hours of motion, ranging from everyday actions to sports and dancing. Motion is standardized to a 22-joint skeleton at 20 frames per second, and the dataset is doubled in size through left-right mirroring. It's the most commonly used training and evaluation benchmark for text-to-motion generation (input a sentence, output a 3D motion clip); motion-diffusion models such as MDM all report results on it. For humanoid robots, one path to making a robot perform actions on command is to first generate human motion from text, then retarget that motion onto the robot.
ExampleGiven the input “a person walks forward and then sits down,” a model trained on HumanML3D generates a 3D skeletal motion of walking followed by sitting, which is then retargeted for a humanoid robot to track and execute.
- Also called
- Text-Annotated 3D Human Motion Dataset
- Related
- Text-to-Motion · AMASS (Archive of Motion Capture as Surface Shapes) · MDM (Motion Diffusion Model) · Motion Retargeting · SMPL · Motion Tracking
- Sources
- HumanML3D GitHub