Embodied AI Glossary中文

Motion Capture

动作捕捉MoCapEssential

Recording the motion of a person or object into a computer with high precision, using cameras or wearable sensors.

Motion capture (mocap) is technology for recording human or object motion into a computer at high precision; it was first used heavily in film, game animation, and sports analysis. There are two main approaches. Optical mocap sticks reflective markers on the body and triangulates their positions with multiple infrared cameras, reaching millimeter-level accuracy or better — Vicon and OptiTrack are leading vendors. Inertial mocap instead straps IMUs (sensors that measure angular velocity and acceleration) to the body, needs no external cameras, and suits outdoor or large-scale use — Xsens is a leading example. There's also markerless mocap, which estimates pose from ordinary video alone. In embodied AI, mocap data, once passed through motion retargeting, can train whole-body motion tracking for humanoid robots, and is also used for teleoperation and to provide ground-truth pose for datasets.

ExampleAMASS unifies 15 optical mocap datasets into the SMPL format — a human body mesh model that represents shape and pose with a small number of parameters — covering more than 300 subjects and over 10,000 motion sequences; it's commonly used to train humanoid robots' motion-tracking policies.

Also called
MoCap
Related
Optical Motion Capture · Inertial Motion Capture · Markerless (Video-Based) Motion Capture · Motion Retargeting · AMASS (Archive of Motion Capture as Surface Shapes) · Data Glove
Sources
Wikipedia: Motion capture
AMASS: Archive of Motion Capture as Surface Shapes

See it in the full glossary →