Real2Render2Real
Real2Render2Real(R2R2R)R2R2RAdvancedA method that renders large amounts of robot training data from a phone scan plus one video of a human demonstration.
Real2Render2Real is a data-generation method released in May 2025 by UC Berkeley (Ken Goldberg's group) and the Toyota Research Institute. It needs only two inputs: a 3D scan of an object taken with a phone, and one video of a person performing the task by hand. The system reconstructs the object's shape with 3D Gaussian splatting, tracks the object's 6-DoF motion throughout the video, converts it into a mesh, pairs it with a robot model, and re-renders thousands of demonstrations with varied object positions and camera viewpoints. It only renders images and does no dynamics simulation at all (collisions are turned off), so there's no need to tune physical parameters or use a real robot. The paper reports that a model trained on data generated from just 1 human demonstration matches the performance of one trained on 150 teleoperated demonstrations, and that generation takes about 1/27th the time of teleoperation.
ExampleScanning a cup with a phone and recording one video of a person placing the cup onto a plate, R2R2R can render thousands of image-action demonstrations of a robot arm performing the same motion, used to train π0-FAST or a diffusion policy.
- Also called
- R2R2R
- Related
- Synthetic Data · 3D Gaussian Splatting · Real-to-Sim-to-Real · Human Video Data · Teleoperation · Diffusion Policy
- Sources
- Real2Render2Real (arXiv:2505.09601)
Real2Render2Real 项目主页 (Chinese) - As of
- 2025-05