Human Pose Estimation
人体姿态估计HPECommonFinding the positions of a person's body joints from an image or video and connecting them into a skeleton.
Human pose estimation takes an image or video as input and outputs the positions of a person's joints — shoulders, elbows, wrists, hips, knees, ankles, and so on — as 2D pixel coordinates or 3D positions, which together form a skeleton. The COCO dataset annotates 17 keypoints per person and scores predictions with object keypoint similarity (OKS), which normalizes by body scale; Carnegie Mellon's OpenPose (CVPR 2017) first finds all the joints in an image and then assigns them to individual people, running in real time even with multiple people in frame; Google's MediaPipe can output 33 body points. For embodied AI, this is the first step in turning human motion into robot data: teleoperation reads the operator's pose in real time, or motion is extracted from human video, and motion retargeting then maps it onto the robot's joints.
ExampleStanford's HumanPlus uses just one RGB camera to estimate the operator's body and hand pose in real time, letting a custom 33-DOF humanoid mimic it synchronously — used both for teleoperation and to collect demonstration data.
- Also called
- HPE, Human Keypoint Detection
- Related
- Keypoint Detection · Markerless Motion Capture · Motion Retargeting · SMPL · Hand Pose Estimation · MediaPipe
- Sources
- Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields (OpenPose, arXiv:1611.08050)
COCO Keypoint Evaluation (OKS)
MediaPipe Pose Landmarker 官方文档 (Chinese)