Embodied AI Glossary中文

Data Flywheel

数据飞轮Essential

A self-reinforcing loop: deployment produces data, the data improves the model, and the better model gets deployed even more widely.

A data flywheel is a self-reinforcing loop: a model is deployed, generates new data through real use — successes, failures, human corrections — and that data, once filtered and labeled, is used to improve the model. A better model can then take on more tasks and reach more deployments, which brings back still more data. NVIDIA defines it as a self-improving loop that keeps refining a model using data collected from its own interactions. The self-driving industry adopted this kind of approach early — Tesla calls its version the data engine. Robot data is expensive to collect, so companies generally want real-world deployment to help spread out that cost; but the flywheel can only start turning once a robot is already useful enough to be deployed in the real world.

ExamplePhysical Intelligence's π*0.6 uses a method called RECAP, which feeds a robot's own autonomous-execution data — from folding laundry, brewing espresso on a professional machine, and assembling boxes in real households — back into training, alongside expert teleoperation corrections. The paper reports throughput on the hardest tasks more than doubling, with the failure rate roughly cut in half.

Also called
Data Closed Loop, Data Engine
Related
Deployment Data Backflow · Human Intervention Data · RECAP · Self-improvement · Human-in-the-Loop · Fleet Learning (Learning While Deploying)
Sources
NVIDIA Glossary: What Is a Data Flywheel?
π*0.6: a VLA That Learns From Experience (arXiv 2511.14759)
As of
2025-11

See it in the full glossary →