Embodied AI Glossary中文

ReKep

Common

Has a large model write task constraints between keypoints as code, then solves for robot motion with optimization.

ReKep (Relational Keypoint Constraints) is a training-free manipulation method proposed in September 2024 by Fei-Fei Li's group at Stanford together with Columbia University. It first uses DINOv2 visual features to find candidate 3D keypoints in an RGB-D image, then gives the numbered, annotated image and a language instruction to GPT-4o, which writes several Python functions: each takes keypoint coordinates as input and outputs a cost, expressing a relation like “the spout must line up with the cup's opening,” with a task optionally broken into multiple stages. A hierarchical optimizer then solves in real time for a sequence of end-effector poses, and replans in closed loop as it tracks the keypoints. This means no task-specific training data is needed to pour tea, fold clothes, or put on shoes with both arms, on either single-arm or bimanual robots — a representative example of the “large model writes constraints, optimizer produces motion” approach.

ExampleGiven the instruction “pour tea into the cup,” the constraints GPT-4o writes include: during the grasp stage, the gripper must be near a keypoint on the pot's handle; during the pour stage, the spout keypoint must sit directly above the cup's opening keypoint with the pot tilted. The optimizer computes the full motion from these.

Also called
Relational Keypoint Constraints, ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
Related
VoxPoser · OmniManip · Affordance · Semantic Keypoints · DINOv2 · LLM-based Task Planning
Sources
ReKep 项目主页 (Chinese)
As of
2024-09

See it in the full glossary →