Embodied AI Glossary中文

Grasp Planning

抓取规划Common

Computing where and how a gripper or dexterous hand should grip an object — position, orientation, and finger configuration.

Grasp planning (also called grasp synthesis) answers ‘where and how to grip’: it outputs a 6D grasp pose (position plus orientation) and opening width for a parallel-jaw gripper, or finger joint angles for a dexterous hand. Motion planning then generates a collision-free path, typically moving to a pre-grasp pose before closing the gripper. Early methods were analytic: given a known object model and friction coefficient, they searched for grasps that maximize quality metrics such as force closure (contact forces able to resist an external force from any direction). A 2014 survey by Bohg et al. grouped data-driven methods into three categories based on whether the object had been seen before: known, similar, or novel. Later, deep learning began predicting grasps directly from depth images or point clouds, as in Dex-Net, Contact-GraspNet, and AnyGrasp. End-to-end vision-language-action (VLA) models usually don't treat this as a separate step, but it remains common in industrial bin-picking and modular pipelines.

ExampleUC Berkeley's Dex-Net 2.0 trained a grasp-quality network called GQ-CNN on 6.7 million synthetic point clouds paired with analytic grasp metrics; on an ABB YuMi robot it planned a grasp in about 0.8 seconds and reached 93% success on 8 known objects.

Also called
Grasp Synthesis, Grasp Pose Planning
Related
Grasping · Grasp Pose Detection · Pre-grasp Pose · Force Closure · Grasp Quality Metric · Motion Planning
Sources
Data-Driven Grasp Synthesis - A Survey (Bohg et al., IEEE T-RO 2014, arXiv 1309.2660)
Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics (arXiv 1703.09312)

See it in the full glossary →