Cold Start
冷启动CommonRunning supervised fine-tuning on a small set of high-quality demonstrations before reinforcement learning, to give the model a decent starting point.
In large-model and VLA training, a cold start means running supervised fine-tuning (SFT — training directly on correct-answer data) on a small amount of high-quality data before reinforcement learning begins, producing an initial model that can already attempt the task reasonably. The term became popular because of DeepSeek-R1 (January 2025): R1-Zero, trained with pure RL, produced hard-to-read output that mixed languages, so the team first fine-tuned on a few thousand long chain-of-thought examples before running RL. This matters because reinforcement learning relies on trial and error, and an initial model that almost never succeeds gets almost no reward signal, making training slow and unstable. The embodied-AI equivalent is training a base policy with behavior cloning on demonstration data first, then fine-tuning it with RL in simulation or on a real robot. Note that “cold start” in recommender systems — a new user with no history — is an unrelated meaning of the same term.
ExampleSimpleVLA-RL (2025) first ran SFT with just one demonstration per task, reaching 17.3% success on LIBERO-Long; using that as the starting point for reinforcement learning then raised it to 91.7%.
- Also called
- Cold-start SFT, Cold-start Data, Cold-start Fine-tuning
- Related
- Supervised Fine-Tuning · Reinforcement Fine-Tuning (RL Fine-Tuning) · Behavior Cloning · Post-training · Sparse Reward · SimpleVLA-RL
- Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (arXiv 2501.12948)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning (arXiv 2509.09674) - As of
- 2025-09