RVT-2
AdvancedNVIDIA's multi-view 3D manipulation policy that learns millimeter-precision insertion from about 10 demonstrations per task.
RVT-2 was released by an NVIDIA research team (Ankit Goyal, Dieter Fox, and others) in June 2024, published at RSS 2024, an upgrade of RVT (Robotic View Transformer). Methods in this family first re-render the point cloud from an RGB-D camera into several virtual-viewpoint images, then use a Transformer on these images to predict the next key pose — where the gripper should go, its orientation, and whether to open or close — before handing off to a motion planner for execution. RVT-2 adds coarse-to-fine multi-stage inference: it first finds a rough region in the whole scene, then zooms in on that region for precise prediction; it also conditions rotation prediction on position, and speeds things up with a custom renderer and a more efficient training implementation. The result is 6x faster training and 2x faster inference than RVT, with multi-task success on RLBench rising from 65% to 82%; on a real robot, using just one RGB-D camera and about 10 demonstrations per task, it can perform millimeter-precision tasks like inserting a peg or plugging in a plug.
ExampleAbout 10 demonstrations of 'insert the pin into the hole' are recorded on a real robot arm, and RVT-2 learns to align and insert into a hole with only a very small clearance.
- Also called
- Robotic View Transformer 2, RVT-2: Learning Precise Manipulation from Few Demonstrations
- Related
- PerAct · Keyframe Action Prediction · Multi-View · RLBench · Peg-in-Hole Insertion · Few-shot
- Sources
- RVT-2: Learning Precise Manipulation from Few Demonstrations (arXiv 2406.08545)
RVT-2 project page - As of
- 2024-06