GVHMR
AdvancedA method that recovers a person’s 3D motion in world coordinates from monocular video.
GVHMR is a method from Zhejiang University’s ZJU3DV team, published at SIGGRAPH Asia 2024: given a monocular video, it outputs SMPL-X human body parameters (a parametric human body model) and the person’s motion trajectory in world coordinates. The difficulty is that the camera itself is also moving — estimating pose only in the camera’s coordinate frame can’t tell which way the person is walking or whether their body is upright. GVHMR defines a “gravity-view” coordinate frame for every frame: one axis aligned with gravity, another referencing the camera’s line of sight, predicts body orientation within this frame, and then converts back to world coordinates using the camera’s relative rotation from visual odometry or a gyroscope. It predicts frame by frame in parallel, avoiding the error accumulation that autoregressive methods like WHAM suffer on long videos. In embodied AI, it is commonly used to extract human motion from internet videos, which is then retargeted onto a humanoid robot.
ExampleThe GMR general-purpose motion retargeting tool supports first extracting human motion from a monocular video with GVHMR, then retargeting it onto a humanoid robot.
- Also called
- GVHMR: World-Grounded Human Motion Recovery via Gravity-View Coordinates
- Related
- Human Mesh Recovery · SMPL · Markerless (Video-Based) Motion Capture · General Motion Retargeting · Motion Retargeting · Visual Odometry
- Sources
- arXiv 2409.06662: World-Grounded Human Motion Recovery via Gravity-View Coordinates
GVHMR 项目主页(ZJU3DV) (Chinese)
GMR: General Motion Retargeting(GitHub README) - As of
- 2024-12