Hand Pose Estimation
手部姿态估计CommonEstimating the positions of a human hand's joints and how the fingers are bent, from an image or sensor data.
Hand pose estimation infers the positions of a human hand's joints and its finger configuration from an image, depth data, or headset sensor data. Two outputs are common: 21 2D or 3D keypoints (the wrist plus 4 points per finger), which is what MediaPipe outputs; or a parametric hand mesh, most commonly the MANO model (proposed in 2017, with 778 vertices controlled by pose and shape parameters), with HaMeR using a large vision transformer to regress a MANO mesh directly from a single image. The difficulty is that fingers are thin and prone to self-occlusion, and are further blocked by whatever object is being held. For embodied AI, this is the first step in turning human hand motion into robot motion: teleoperation gets real-time keypoints from a headset's hand tracking, which motion retargeting then converts to drive a dexterous hand; learning from human video likewise requires the hand's trajectory to be estimated first, as a pseudo-action label.
ExampleOpen-TeleVision uses an Apple Vision Pro to get the operator's hand keypoints in real time, which dex-retargeting then optimizes into dexterous-hand joint angles; when the operator makes a fist, the Unitree H1's hand follows suit.
- Also called
- Hand Tracking, Hand Mesh Recovery
- Related
- MANO · HaMeR · MediaPipe · Motion Retargeting · Hand-Object Interaction · Human Pose Estimation
- Sources
- HaMeR: Reconstructing Hands in 3D with Transformers (arXiv:2312.05251)
MANO 官网 (Chinese)
Open-TeleVision (arXiv:2407.01512)