Embodied AI Glossary中文

OMOMO

OMOMO 人-物交互动作数据集Advanced

About 10 hours of optical motion-capture data from Stanford, recording full-body motion while carrying everyday objects.

The OMOMO dataset comes from a SIGGRAPH Asia 2023 paper by Jiaman Li, Jiajun Wu, and C. Karen Liu at Stanford. The authors used a 12-camera Vicon optical motion-capture rig to record 17 subjects carrying and dragging 15 everyday objects (a mop, a floor lamp, a chair, a table, boxes, and more), capturing full-body motion for about 10 hours in total, along with each object's 3D geometry, its own motion, and the human motion in SMPL-X format. The paper's method takes only the object's motion as input, and uses a conditional diffusion model to first predict hand positions, then generate full-body pose. This is one of the few datasets capturing full-body interaction with large objects, and it's now commonly used as reference motion for teaching humanoid robots whole-body manipulation like carrying, pushing, and pulling — OmniRetarget, for example, retargets it onto the Unitree G1.

ExampleGiven the motion trajectory of a chair being dragged, the OMOMO method generates a full-body human motion of bending down to grab the chair back and dragging it while walking.

Also called
OMOMO Dataset, Object Motion Guided Human Motion Synthesis
Related
Human-Object Interaction · Optical Motion Capture · SMPL · Diffusion Model · OmniRetarget · Motion Retargeting
Sources
Object Motion Guided Human Motion Synthesis (arXiv 2309.16237)
Object Motion Guided Human Motion Synthesis (arXiv HTML 全文) (Chinese)
As of
2023-09

See it in the full glossary →