Co-training
协同训练CommonMixing a small amount of target-robot data with data from other sources in fixed proportions to train one model.
In embodied AI, co-training means mixing a small amount of target-robot data together with data from other sources, in a set ratio, into the same training batches for one shared model. The other sources can be web image-text QA data, data from other robots, or simulation data. Google's RT-2 (2023) fine-tuned on robot trajectories together with web tasks like visual question answering, calling it co-fine-tuning; π0.5 (2025) trains jointly on multi-robot data, high-level semantic predictions, and web data. The point is to work around having too little real-robot data: mixing in other data preserves visual and language common sense and improves generalization, and the key is getting the mixing ratio right. Note this is a different sense from the “co-training” Blum and Mitchell proposed in 1998, which is a semi-supervised learning method.
ExampleMobile ALOHA had only 50 mobile-manipulation demonstrations per task; co-training with existing static ALOHA bimanual data raised success rates by up to 90 percentage points. In a 2025 sim-and-real co-training study from NVIDIA and collaborators, mixing in simulation data raised real-robot task success rates by 38% on average.
- Also called
- Joint Training, Mixed Training, Co-fine-tuning, Sim-and-Real Co-training
- Related
- Data Mixture · Sim-to-Real Transfer · Cross-Embodiment Data · RT-2 · π0.5 · Mobile ALOHA
- Sources
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (arXiv 2307.15818)
Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation (arXiv 2401.02117)
Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation (arXiv 2503.24361) - As of
- 2025-04