Self-improvement
自我提升AdvancedA robot retrains itself on data from its own practice, getting steadily better with less reliance on human data.
Self-improvement refers to a loop in which a model, after starting from a small amount of human data, generates new data by interacting with the environment on its own, automatically judges whether it did well, and retrains on that, aiming to stop data growth from depending entirely on human demonstrations. The key is automatically judging success or failure: common tools are success detectors, reward models, or value functions, paired with filtered behavior cloning or reinforcement learning to update the policy. Representative work includes DeepMind's 2023 RoboCat, which trains its next generation on data from its own practice, and Ghasemipour and colleagues' 2025 Self-Improving Embodied Foundation Models, which has the model predict how many steps remain, yielding both a reward and a success detector at once so a fleet of robots can practice autonomously, more sample-efficient than simply collecting more demonstrations. π*0.6, which learns from autonomous experience and human corrections using RECAP, belongs to this category too.
ExampleOne round of RoboCat's loop: humans teleoperate 100 to 1,000 new demonstrations for a task, a specialist branch model is fine-tuned from them, the branch practices autonomously about 10,000 times on average, and the new and old data are merged to train the next RoboCat generation.
- Also called
- Self-Improving, Autonomous Improvement
- Related
- RoboCat · Data Flywheel · Success Detector · Real-World Reinforcement Learning · π*0.6 · Rejection Sampling Fine-Tuning
- Sources
- Google DeepMind Blog: RoboCat: A self-improving robotic agent
Bousmalis et al. 2023: RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation
Ghasemipour et al. 2025: Self-Improving Embodied Foundation Models - As of
- 2025-09