OKAMI
AdvancedA 2024 UT Austin and NVIDIA method that teaches a humanoid robot to manipulate objects from watching a single human video.
OKAMI was proposed by Yuke Zhu's group at UT Austin together with NVIDIA Research, posted to arXiv in October 2024, an oral presentation at CoRL 2024. The goal is to have a humanoid robot learn a manipulation task from watching just a single RGB-D video of a human demonstration, with no teleoperation data collection needed. The first step analyzes the video: GPT-4V identifies task-relevant objects, Grounded-SAM segments and Cutie tracks them, and the person's body and hand motion is reconstructed (using the SMPL-H model) to get a reference plan. The second step is object-aware motion retargeting: it first locates where the objects are in the current scene, then adjusts the human arm trajectory to that object position before mapping it onto the robot, with finger motion transferred along with it. The experimental platform is a Fourier GR1 humanoid fitted with two 6-degree-of-freedom Inspire dexterous hands. Successfully executed trajectories can also serve as data for training a closed-loop visuomotor policy.
ExampleAcross 6 tasks — bagging groceries, sprinkling salt, putting a snack on a plate, closing a laptop, and others — OKAMI's average success rate is 71.7%, 58.3 percentage points above the baseline ORION; a visuomotor policy trained on the trajectories it produces reaches an average success rate of 79.2%.
- Also called
- Teaching Humanoid Robots Manipulation Skills through Single Video Imitation
- Related
- Imitation from Observation · Motion Retargeting · Human Video Data · One-shot Imitation Learning · Humanoid Robot · Fourier GR-1
- Sources
- OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation (arXiv 2410.11792)
OKAMI 项目页 (Chinese) - As of
- 2024-11