ReKep
CommonHas a large model write task constraints between keypoints as code, then solves for robot motion with optimization.
ReKep (Relational Keypoint Constraints) is a training-free manipulation method proposed in September 2024 by Fei-Fei Li's group at Stanford together with Columbia University. It first uses DINOv2 visual features to find candidate 3D keypoints in an RGB-D image, then gives the numbered, annotated image and a language instruction to GPT-4o, which writes several Python functions: each takes keypoint coordinates as input and outputs a cost, expressing a relation like “the spout must line up with the cup's opening,” with a task optionally broken into multiple stages. A hierarchical optimizer then solves in real time for a sequence of end-effector poses, and replans in closed loop as it tracks the keypoints. This means no task-specific training data is needed to pour tea, fold clothes, or put on shoes with both arms, on either single-arm or bimanual robots — a representative example of the “large model writes constraints, optimizer produces motion” approach.
ExampleGiven the instruction “pour tea into the cup,” the constraints GPT-4o writes include: during the grasp stage, the gripper must be near a keypoint on the pot's handle; during the pour stage, the spout keypoint must sit directly above the cup's opening keypoint with the pot tilted. The optimizer computes the full motion from these.
- Also called
- Relational Keypoint Constraints, ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
- Related
- VoxPoser · OmniManip · Affordance · Semantic Keypoints · DINOv2 · LLM-based Task Planning
- Sources
- ReKep 项目主页 (Chinese)
- As of
- 2024-09