Keypoint Detection
关键点检测CommonLocating a handful of pre-defined, meaningful points in an image, like a wrist, a cup's handle, or a box's corner.
Keypoint detection locates a small number of points with predefined meaning in an image, such as a wrist, a cup's handle, or a box's corner. It's more precise than a bounding box and lighter-weight than a segmentation mask; it also differs from feature points like SIFT or ORB, which only need to be locally recognizable and carry no fixed meaning. The most common case is human body keypoints: COCO labels each point as ‘not labeled,’ ‘occluded,’ or ‘visible,’ and scores predictions with OKS similarity. In robot manipulation, MIT's kPAM (2019) represents a whole category of objects with a few 3D semantic keypoints, so swapping in a differently shaped cup still works under the same rule — handling shape variation within a category better than estimating a single 6D pose would.
ExampleTo hang various mugs on a mug rack, the robot first detects each mug's base, rim, and handle as 3D keypoints, then plans motion so the handle lines up with the hook — the same rule works even though the mugs differ in size and shape.
- Also called
- Keypoints, Keypoint Localization
- Related
- Human Pose Estimation · Semantic Keypoints · Feature Points · 6D Object Pose Estimation · ReKep · Tracking Any Point
- Sources
- COCO Data Format(keypoints 标注格式) (Chinese)
kPAM: KeyPoint Affordances for Category-Level Robotic Manipulation (arXiv:1903.06684)