Semantic Keypoints
语义关键点AdvancedRepresenting an object with a handful of meaningful points on it — a mug’s handle, a kettle’s spout — to make planning manipulation easier.
Semantic keypoints are a small number of 3D points on an object that carry clear meaning — for example, the center of a mug’s handle, the center of its base, or the heel of a shoe. Unlike 6D pose, this does not require an exact CAD template for every object; objects within the same category whose shapes vary a lot can still have corresponding points found on them, which makes the representation well suited to category-level generalization. kPAM, from Russ Tedrake’s group at MIT in 2019, applied this to robot manipulation: it first detects semantic 3D keypoints, then expresses the task as geometric constraints on those points (such as “hang the mug’s handle on the hook” or “press the mug’s base flat against the table”), and solves an optimization for the arm’s target pose, letting it manipulate new objects it has never seen. Recent work such as ReKep and OmniManip has vision-language models propose keypoints and constraints directly from an image, which is why the term “task keypoints” is also common.
ExampleIn kPAM, detecting only a few keypoints — the handle and the base — on mugs of many different, unseen shapes is enough to plan the motion for hanging each one on a mug rack.
- Also called
- Task Keypoints, Semantic 3D Keypoints
- Related
- Keypoint Detection · ReKep · OmniManip · 6D Object Pose Estimation · Category-Level Pose Estimation · Affordance
- Sources
- arXiv 1903.06684: kPAM: KeyPoint Affordances for Category-Level Robotic Manipulation