Real-Robot-Data Camp vs. Sim-Data Camp
真机派 / 仿真派CommonThe debate over whether robot training data should mainly come from real robots or from simulation.
This is how China's robotics industry refers to its ongoing debate over the data pipeline for embodied AI. The real-data camp argues that the fine details of real physical interaction are too hard for simulation to reproduce faithfully, and favors collecting large volumes of real-robot data via teleoperation — Physical Intelligence and AgiBot's AgiBot World dataset are representative examples. The sim-data camp argues real-robot collection is too slow and expensive, and favors generating synthetic data at scale in simulators using domain randomization and procedural generation, then applying sim-to-real transfer — Galbot's use of synthetic data to train grasping models, and NVIDIA's simulation toolchain, are representative examples. In practice most teams blend both approaches and are increasingly adding human-video data too; the “data pyramid” concept is exactly this idea of layering different data types by how much of each is used.
ExampleGalbot's GraspVLA is pretrained mainly on roughly a billion synthetic grasping examples, while AgiBot World is a large-scale dataset collected from real robots at a dedicated data-collection facility.
- Also called
- Real-Data Camp / Sim-Data Camp, Real-World vs. Simulation Data Debate
- Related
- Real-Robot Data · Simulation Data · Synthetic Data · Sim-to-Real Transfer · Data Pyramid · Technical Route Debate
- Sources
- AgiBot World (GitHub)
NVIDIA Isaac Lab - As of
- 2025