Embodied AI Glossary中文

Markerless Motion Capture

无标记动捕Common

Recovering 3D human motion straight from ordinary video, with no reflective markers worn on the body.

Traditional optical motion capture sticks reflective markers on the body and tracks them with multiple infrared cameras, giving high precision but requiring expensive equipment that only works in a dedicated studio space. Markerless motion capture removes the markers, instead relying on human pose estimation to find joints from one or a few ordinary camera videos, then fitting the result to a 3D skeleton or a parametric model such as SMPL. Stanford's OpenCap, from 2023, films with two or more iPhones and reports an average joint-angle error of about 4.5°; Zhejiang University's GVHMR (SIGGRAPH Asia 2024) can recover human motion in world-frame coordinates from a single monocular video alone. Markerless capture still lags behind optical capture on occlusion, fast motion, and finger detail. For embodied AI, it makes it possible to turn large amounts of human video into motion references for humanoid robots.

ExampleFilming a dance with a phone, then recovering it into a world-frame SMPL motion sequence with GVHMR and retargeting it onto a humanoid's joints, gives a reference motion for reinforcement-learning motion tracking.

Also called
Video-Based Motion Capture
Related
Motion Capture · Optical Motion Capture · Human Pose Estimation · Human Mesh Recovery · SMPL · Motion Retargeting
Sources
OpenCap: Human movement dynamics from smartphone videos (PLOS Computational Biology, 2023)
GVHMR: World-Grounded Human Motion Recovery via Gravity-View Coordinates (arXiv:2409.06662)

See it in the full glossary →