10Training & Learning Methods
How models are actually trained: imitation learning, reinforcement learning, pretraining and fine-tuning, plus tricks that make it more stable. · 203 terms
- 10.1Overview of learning paradigms5
- 10.2Training fundamentals25
- 10.3Loss functions and training objectives13
- 10.4Imitation learning7
- 10.5Core reinforcement-learning concepts20
- 10.6Classic reinforcement-learning algorithms22
- 10.7Where rewards come from17
- 10.8Offline reinforcement learning13
- 10.9Pretraining and representation learning24
- 10.10Fine-tuning and post-training25
- 10.11From simulation to real robots23
- 10.12Inference-time gains and fast adaptation9
10.1Overview of learning paradigms
With data ready, meet the basic ways to learn: with labels, without them, self-generated labels, from demonstration, or from reward.
- Training a model on input–correct-answer pairs so it learns to predict the answer from the input alone.
- Using only unlabeled data and letting the model discover structure in it on its own.
- Training a model on “questions and answers” constructed from the data itself, with no human labeling needed.
- Teaching a robot a skill by having it imitate an expert's — usually a human's — demonstrated actions.
- A machine-learning approach where an agent improves its behavior through repeated trial and error, guided by a reward signal.
10.2Training fundamentals
Whatever the method, training means computing loss, taking gradients, and updating parameters, while guarding against overfitting and instability.
- Training / Validation / Test Set训练集 / 验证集 / 测试集Splitting a dataset into three parts: one to learn from, one to tune and select a model, and one saved for a final, honest check.
- The real data treated as the “correct answer” for training and evaluating a model.
- Loss Function损失函数A single number measuring how far a model's prediction is from the correct answer; training just means shrinking it.
- Gradient Descent梯度下降An optimization method that nudges parameters step by step opposite the loss function's gradient, making the loss smaller each time.
- Backpropagation反向传播The chain-rule algorithm that computes gradients layer by layer from the output backward; the core algorithm for training neural networks.
- Optimizer优化器The algorithm that turns gradients into parameter updates during training, such as SGD, Adam, or AdamW.
- AdamWAdamW 优化器A version of the Adam optimizer that separates weight decay from the gradient update; the default choice for training most large models.
- A training setting chosen by a person beforehand rather than learned by the model itself, like the learning rate or batch size.
- The size of the step taken at each parameter update in gradient descent; one of the most important hyperparameters.
- Learning Rate Schedule (Warmup and Cosine Decay)学习率调度(预热与余弦退火)Changing the learning rate over the course of training: ramping it up during warmup, then easing it down along a cosine curve.
- Batch Size批大小The number of samples fed through the model together for each single parameter update.
- Epoch训练轮次One complete pass of the model through the entire training set.
- The point in training where the loss or return stops changing much, meaning the model has settled into a stable state.
- Checkpoint检查点A saved snapshot of a model's parameters, taken during or after training, that can be loaded to run inference or resume training.
- Alchemy (Deep-Learning Slang)炼丹 / 调参(黑话)Chinese deep-learning slang for training models and tuning hyperparameters by trial and error, likening it to alchemy.
- Random Seed and Reproducibility随机种子与可复现性Fixing the random number generator's starting point so runs can be repeated, and checking results across several seeds to rule out luck.
- Overfitting过拟合When a model memorizes the training data too closely, so its performance drops on new data it hasn't seen.
- Underfitting欠拟合A model too weak or too undertrained to learn even the patterns already present in the training data.
- Adding constraints or penalties during training to keep a model from memorizing training data, improving performance on new data.
- Dropout随机失活Randomly “turning off” some neurons during training, a regularization technique that keeps a network from memorizing the training data.
- Stopping training once validation performance stops improving, and keeping the checkpoint that performed best.
- Vanishing / Exploding Gradients梯度消失 / 梯度爆炸Gradients shrinking to near zero or blowing up as they're multiplied across many layers during backpropagation, stalling deep-network training.
- Scaling down an overly large gradient to stay within a threshold, so one bad update can't wreck the model.
- Stop-Gradient梯度阻断A value is used normally in the forward pass but treated as a constant during backpropagation, blocking its gradient.
- A weighted average of past values that decays exponentially, often used to get smoother, more stable model weights.
10.3Loss functions and training objectives
Expanding on loss functions: the different objectives used for regression, classification, sequence prediction, and diffusion generation.
- The average of the squared difference between predictions and ground truth; the most common loss for regression.
- L1 LossL1 损失A loss function that measures error as the absolute value of the difference between a prediction and the ground truth.
- A measure of how far a model's predicted probability distribution is from the correct answer; the standard loss for classification.
- Maximum Likelihood Estimation最大似然估计 / 负对数似然Finding the parameters that make the observed data most probable; taking the negative log of that probability gives a loss to minimize.
- A measure of how far one probability distribution is from another; not symmetric between the two.
- Next-Token Prediction下一个 token 预测The training objective of predicting the next token in a sequence given everything that came before it.
- Teacher Forcing教师强制Training a sequence model by feeding it the real previous step at every step, instead of its own prediction.
- Self Forcing自强制Training a video model by having it keep generating from its own previously generated frames, removing the train-test mismatch.
- A tractable lower bound on the log-likelihood of data; VAEs and diffusion models both train by maximizing it.
- A diffusion model's training objective: add noise to clean data and have the network predict exactly the noise that was added.
- Score Matching分数匹配Training a network to estimate the gradient of a data distribution's log-density, without needing its normalizing constant.
- Flow Matching Loss流匹配损失Training a network with mean squared error to predict the velocity that points from noise toward the real data.
- Auxiliary Loss / Auxiliary Task辅助损失 / 辅助任务An extra prediction task and loss term added alongside the main training objective, to help the model learn better features.
10.4Imitation learning
Applying supervised training to learn actions from human demonstrations, and why it tends to drift further off course over time.
- Behavior Cloning行为克隆Treating expert demonstrations as labeled data and using supervised learning to directly map observations to actions.
- Small per-step deviations in an imitation-learned policy add up, pushing the robot into unfamiliar states where it fails.
- DAggerDAgger(数据集聚合)Letting the policy run itself, having an expert label the correct action at the states it actually reaches, then retraining on the combined data.
- Imitation learning latching onto a cue correlated with the expert's actions but not actually its cause, then failing at deployment.
- Action / State Normalization动作与状态归一化Scaling each action and state dimension to a common range before training, then converting model output back to real units at inference.
- Behavior cloning that feeds the policy both the current observation and the goal it's meant to reach.
- Imitation from Observation从观测中模仿学习Learning to imitate from only the demonstrator's states or video, with no recorded action labels at all.
10.5Core reinforcement-learning concepts
The second path besides demonstration is trial and error by reward: first, the basics of reward, return, and value.
- Markov Decision Process马尔可夫决策过程The standard mathematical framework for an agent's loop of seeing a state, acting, getting a reward, and moving to a new state.
- Reward Function奖励函数The scoring rule in reinforcement learning that rates an agent's behavior at each step, defining what counts as doing well.
- Return回报The sum of all rewards from a given moment onward, usually discounted so that later rewards count for less.
- Discount Factor折扣因子The coefficient in reinforcement learning that discounts future rewards, controlling how much the agent weighs long-term payoff.
- The trade-off between trying new actions to find something better and using the best-known action to collect reward now.
- Figuring out, once a final reward or penalty arrives, which of the earlier actions deserve the credit or the blame.
- Value Function价值函数A function estimating how much total reward will follow from a given state if the agent keeps following a given policy.
- Bellman Equation贝尔曼方程A recursive formula that breaks a state's value into the immediate reward plus the discounted value of the next state.
- Q-FunctionQ 函数The expected return of taking a specific action in a state and then following the policy from then on.
- A measure of how much better an action is than the policy's average, equal to the Q-value minus the value function.
- Policy Iteration / Value Iteration策略迭代 / 价值迭代Two classic dynamic-programming algorithms that repeatedly update values or the policy to solve for an optimal policy, given a known model.
- Monte Carlo Methods蒙特卡洛方法(蒙特卡洛回报)Estimating a value by averaging over many random samples; in RL, using a whole episode's actual return to estimate value.
- Updating a value estimate using “this step's reward plus the estimated value of the next state,” without waiting for the episode to end.
- Updating a state's estimated value using the agent's own current estimate of the next state's value.
- How much data or environment interaction is needed to reach a given performance level; using less means being more efficient.
- Reinforcement learning that skips building an environment model and learns a policy or value function directly from trial-and-error data.
- Model-Based Reinforcement Learning基于模型的强化学习Learning a model that predicts how the environment will respond, then using it to plan or “imagine” training data for a policy.
- Learning a world model first, then training the policy inside trajectories the model “imagines”.
- On-Policy同策略Updating only with data the current policy just collected itself, and discarding old data once it's used.
- Off-Policy异策略Reinforcement learning that can use data collected by a different policy, not only data the current policy gathered itself.
10.6Classic reinforcement-learning algorithms
Turning those concepts into concrete algorithms: from policy gradients and PPO to DQN and SAC.
- Policy Gradient策略梯度Computing the gradient of expected return with respect to the policy's parameters directly, and improving the policy along that gradient.
- REINFORCEREINFORCE 算法The earliest policy-gradient algorithm: raise the probability of actions that led to high return across a full trajectory.
- Estimating the advantage function by summing multi-step temporal-difference errors with exponentially decaying λ weights.
- Importance Sampling重要性采样Estimating an expectation under one distribution using samples drawn from another, weighted by the ratio of the two probabilities.
- A policy-gradient algorithm that bounds the KL divergence between old and new policy at every update; PPO's predecessor.
- OpenAI's 2017 reinforcement-learning algorithm that limits how much each policy update can change the policy, making training stable.
- A reinforcement-learning algorithm that samples a group of outputs for the same input and uses their relative scores instead of a value network.
- Q-LearningQ 学习A classic model-free reinforcement-learning algorithm that repeatedly corrects its Q-values toward the immediate reward plus the next state's best Q-value.
- Deep Q-Network深度 Q 网络A reinforcement-learning algorithm that uses a deep neural network to estimate each action's long-term value, its Q-value.
- Storing an agent's past interactions in a buffer and sampling them randomly during training, instead of using each one only once.
- Target Network目标网络A slowly updated copy of the main network, used only to compute training targets, making Q-learning more stable.
- Overestimation BiasQ 值高估Taking a max over noisy Q-value estimates systematically inflates them, biasing an agent toward overrated actions.
- A value function that predicts the full probability distribution of future return, not just its average.
- Deterministic vs. Stochastic Policy确定性策略 / 随机策略A deterministic policy always gives the same action for the same state; a stochastic policy gives a probability distribution to sample from.
- Deep Deterministic Policy Gradient深度确定性策略梯度An actor-critic algorithm that extends DQN to continuous actions, with the policy outputting one deterministic action directly.
- Twin Delayed DDPG双延迟深度确定性策略梯度A continuous-action reinforcement-learning algorithm that adds three fixes to DDPG specifically to curb Q-value overestimation.
- Soft Actor-Critic软演员-评论家An off-policy reinforcement-learning algorithm that pursues high return while also encouraging the policy to stay somewhat random.
- Adding a bonus for policy entropy to the training objective, to encourage the policy to stay random and keep exploring.
- Entropy Collapse / Mode Collapse熵坍缩 / 模式坍缩A model's output diversity collapsing, so it only ever produces a handful of answers or actions.
- Update-to-Data Ratio更新-数据比How many gradient updates are done per step of environment data collected; higher saves data but costs more compute.
- Reinforcement learning where multiple agents learn at once in a shared environment, cooperating or competing.
- Self-Play自博弈Having an agent compete against itself, or past versions of itself, to keep improving through the outcomes.
10.7Where rewards come from
With algorithms in hand, you still need the right reward: hand-written, learned from demonstrations or a model, and what to do when it’s sparse.
- Sparse Reward稀疏奖励A reward given only at a few key moments, such as task completion, with zero reward the rest of the time.
- Dense Reward稠密奖励A reward design that gives informative feedback at nearly every step, showing whether the agent is getting closer to or further from the goal.
- Reward Shaping奖励塑形Adding intermediate guidance rewards on top of a task's original reward, so the agent learns the task faster.
- Designing and debugging a reward function for reinforcement learning so the robot actually learns the intended behavior.
- Reward Hacking奖励黑客An agent finding a loophole in the reward function to score high without actually accomplishing what the designer intended.
- Working backward from an expert's demonstrated behavior to infer the reward function it's implicitly optimizing.
- Using a discriminator to tell expert actions from policy actions, forcing the policy to act more and more like the expert.
- Using a discriminator that judges how much motion resembles motion-capture data, and turning that resemblance into a reward for natural movement.
- Reward Model奖励模型A trained scoring network that outputs a reward telling how good a given behavior or outcome is.
- Progress Reward Model进度奖励模型A model that looks at the current frame and estimates how much of a task is done, used as a reward.
- Success Detector成功检测器A model that judges whether a robot completed a task in a given episode, often used to give reward in RL.
- VLM-as-RewardVLM 作奖励模型Having a vision-language model watch footage against a task description and score the robot's performance as a reward.
- A reward the agent generates for itself from novel or hard-to-predict states, used to drive exploration.
- Unsupervised Skill Discovery无监督技能发现With no task reward, letting an agent practice into a set of distinguishable skills it can call on later.
- Reinforcement learning where both the policy and the value function take the goal as input, letting one model reach many goals.
- Hindsight Experience Replay后见之明经验回放Relabeling a failed attempt as having succeeded at whatever goal it actually reached, so even sparse rewards can be learned from.
- Reinforcement learning split into layers: a high level sets sub-goals, and a low level executes the concrete actions to reach them.
10.8Offline reinforcement learning
Earlier algorithms all learn through live interaction; here a policy learns from a fixed dataset instead, then fine-tunes online.
- Learning while interacting with the environment, continuously collecting new data with the latest policy during training.
- Training a policy using only a fixed, previously collected dataset, with no further interaction with the environment during training.
- The error in offline reinforcement learning that comes from a Q-network guessing wildly at the value of actions absent from the data.
- In offline RL, keeping the new policy's actions close to what the behavior policy in the dataset actually did.
- Conservative Q-Learning保守 Q 学习Deliberately pushing down the Q-values of actions absent from the dataset, so offline reinforcement learning isn't misled by inflated estimates.
- Implicit Q-Learning隐式 Q 学习An offline reinforcement-learning algorithm that only ever values actions already in the dataset, never querying an action it hasn't seen.
- Doing imitation learning weighted by each action's advantage, so higher-advantage actions get imitated more.
- Return Conditioning回报条件化Feeding a policy the return you want it to achieve, so it acts toward that target return.
- Telling the policy how “good” each action was during training, then at deployment asking it only to generate “good” actions.
- Pretraining with offline reinforcement learning on existing data, then letting the robot keep improving through live interaction.
- Calibrated Q-Learning校准 Q 学习Adding a “calibration” constraint to conservative Q-learning, so offline pretraining can transition smoothly into online fine-tuning.
- Reinforcement Learning with Prior Data利用先验数据的强化学习A simple, efficient way to do online RL by mixing offline data half-and-half with freshly collected data in every batch.
- Q-Chunking动作分块强化学习Doing reinforcement learning where both the policy and the Q-function operate on a whole chunk of actions at once.
10.9Pretraining and representation learning
Shifting from reinforcement learning to the large-model approach: pretraining on massive, varied data to learn general-purpose representations.
- Pre-training预训练Training a foundation model on massive general-purpose data first, before adapting it to any specific task.
- Downstream Task下游任务The specific task a pretrained model is ultimately meant to solve, usually reached by further adaptation or fine-tuning.
- Using knowledge learned on one task or domain to help learn another, related task.
- Training a model directly on the target data with randomly initialized parameters, borrowing no pretrained weights at all.
- Scaling Law缩放定律The empirical pattern that model performance improves smoothly, as a power law, with more parameters, data, and compute.
- Ossification骨化(模型骨化)A model's weights seem to “freeze up,” unable to absorb new information even as more training data is added.
- Distributed Training分布式训练(数据并行 / 模型并行)Splitting a single training run across multiple GPUs or machines so larger models and more data can be trained.
- Mixed-Precision Training混合精度训练Running most computation in 16-bit floating point and keeping only the sensitive parts in 32-bit, to save memory and gain speed.
- Accumulating gradients over several small batches before applying one combined parameter update, simulating a larger batch size.
- Storing fewer intermediate activations during the forward pass and recomputing them during backward, trading compute for memory.
- Letting a model automatically learn useful feature vectors from raw data, instead of hand-designing the features.
- Learning features from unlabeled data by pulling similar samples' representations together and pushing dissimilar ones apart.
- InfoNCE LossInfoNCE 损失A contrastive-learning loss that trains a model to pick out the one true positive among many candidates.
- Self-supervised video representation learning: pull same-moment multi-view frames together, push nearby-but-different moments apart.
- Masked Autoencoder掩码自编码器Hiding most patches of an image and training a model to reconstruct them, as self-supervised visual pretraining.
- Encoding raw tactile-sensor readings into general-purpose features that downstream tasks like slip detection can reuse.
- Training a model so its intermediate features match those of an existing pretrained encoder.
- Mapping features from a new modality, like images, into a representation space the language model can already understand.
- Pretraining on Human Videos人类视频预训练Pretraining a robot model on large amounts of video of humans doing things, then fine-tuning it with a small amount of robot data.
- Latent Action Pretraining潜在动作预训练Extracting unlabeled “latent actions” from video, then pretraining a robot model to predict them.
- Multi-Task Learning多任务学习Training one model on several related tasks at once, so it shares knowledge and each task helps the others.
- Co-training协同训练Mixing a small amount of target-robot data with data from other sources in fixed proportions to train one model.
- Positive / Negative Transfer正迁移 / 负迁移When adding other tasks or data makes the target task better, that's positive transfer; when it makes it worse, that's negative transfer.
- Mid-training中训练An extra training stage between pretraining and post-training that uses more targeted data to fill in specific abilities.
10.10Fine-tuning and post-training
After pretraining, adapting to specific tasks: fine-tuning, supervised and RL post-training, plus preventing forgetting and distillation.
- The training stage that follows pretraining, using curated data or reinforcement learning to turn a foundation model into something actually useful.
- Continuing to train an already-trained model on a small amount of new-task data, so it gets better at that specific job.
- Full Fine-Tuning全参数微调Fine-tuning that updates every parameter of a pretrained model, rather than training just a small subset of them.
- Backbone Freezing冻结骨干网络Keeping a pretrained backbone network's parameters fixed during training and updating only the newly added parts.
- Freezing most of a large model's parameters and training only a small added or selected subset for a new task.
- LoRA低秩适配Freezing a large model's original weights and fine-tuning it by training only two small matrices inserted alongside them.
- Adapter适配器Small modules inserted between a frozen large model's layers, with only those small modules trained to adapt it to a new task.
- Prompt Tuning / Soft Prompt提示微调 / 软提示Freezing the whole model and training only a short learnable vector prepended to the input to adapt it to a task.
- Continuing supervised training on a pretrained model using paired input–correct-answer data, to teach it a specific behavior.
- Fine-tuning a model on many tasks written as instruction-and-answer pairs so it learns to follow instructions.
- Rejection Sampling Fine-Tuning拒绝采样微调 / 过滤式行为克隆Letting a model attempt a task many times, keeping only the successful or high-scoring results, and training on those.
- Cold Start冷启动Running supervised fine-tuning on a small set of high-quality demonstrations before reinforcement learning, to give the model a decent starting point.
- Further improving a pretrained or supervised-fine-tuned model with reinforcement learning driven by a reward signal.
- Reinforcement Learning from Human Feedback基于人类反馈的强化学习Having people compare pairs of model outputs to train a reward model, then using reinforcement learning to optimize toward human preference.
- Aligning a model directly from paired “better/worse” samples, without training a reward model or running reinforcement learning.
- Reinforcement Learning with Verifiable Rewards基于可验证奖励的强化学习Training a model with reinforcement learning using rewards a program can automatically check as correct or wrong.
- KL RegularizationKL 正则化Adding a KL-divergence penalty to a training objective so the new policy doesn't drift too far from a reference.
- A neural network's sharp loss, or outright loss, of old abilities after it learns something new.
- Letting a model learn a sequence of new tasks and new data without forgetting what it already learned.
- A VLA training trick that blocks gradients from the action expert back into the VLM backbone, protecting its pretrained knowledge.
- Model Merging模型合并Combining several related models' parameters by weighted averaging, with no extra training required.
- Training a small student model to mimic a large teacher model's outputs, compressing the teacher's ability into a smaller model.
- Having a student policy imitate one or more already-trained teacher policies, compressing or merging their skills.
- On-Policy Distillation在线策略蒸馏Distillation where the student generates its own trajectories, and the teacher scores and corrects it at every step.
- Compressing a diffusion model that needs dozens to hundreds of denoising steps into a student that produces results in one or a few.
10.11From simulation to real robots
Onto real robots: large-scale training in sim, transferring to hardware, then continuing to learn there through trial and human correction.
- Running thousands of simulated environments at once on a single GPU to collect data, training a locomotion policy in just minutes.
- Actor-Learner Architecture (Distributed RL)Actor-Learner 分离架构Splitting “interacting with the environment to collect data” and “updating the network's parameters” across separate processes or machines, run in parallel.
- Population-Based Training基于群体的训练Training a whole population of models at once, periodically copying good weights over bad ones and perturbing hyperparameters.
- A training strategy that starts a model on easy samples or tasks and gradually increases the difficulty.
- In legged reinforcement learning, a training schedule that gradually moves a robot onto harder terrain as it improves.
- Ending an episode immediately and resetting the environment as soon as a failure state, like falling, occurs during training.
- In motion-imitation training, starting each episode from a random point in the reference motion instead of always frame one.
- Applying random transformations to existing training data to create new samples, without collecting anything new.
- Symmetry Augmentation对称性增强(镜像损失)Using a robot's left-right symmetry, by mirroring data or adding a loss term, to make its learned motion symmetric.
- Domain Adaptation领域自适应Making a model trained on one data distribution (the source domain) work well on a different distribution (the target domain).
- Extra information available only during training, not at deployment, like a simulator's exact terrain shape or friction values.
- Asymmetric Actor-Critic非对称演员-评论家Letting the critic see the full simulated state while the actor sees only what a real robot could actually observe.
- Teacher-Student Distillation教师-学生蒸馏Training a teacher policy that can see privileged information first, then having a student that only uses real sensors imitate it.
- Letting a robot learn by trial and error directly in the real environment, rather than only training in simulation.
- Letting a robot keep learning through continuous interaction, without a person resetting the environment every episode.
- Maximizing reward while guaranteeing safety constraints are not violated, during both training and deployment.
- Keeping an existing base controller and using reinforcement learning to learn only a correction on top of its output.
- Noise-Space Policy Steering噪声空间策略引导Keeping a diffusion policy's weights fixed and using reinforcement learning to pick its input noise instead, changing its output.
- Keeping a person inside a robot's training or operating loop to correct, take over, or give feedback in real time.
- Imitation learning where a human gives feedback while the robot acts, and the policy improves from it online.
- Human-Gated DAgger人工门控 DAggerA DAgger variant where a human takes over just before the robot is about to err, and only that takeover data gets used for retraining.
- Fleet Learning (Learning While Deploying)机群学习 / 部署中学习Having a whole fleet of already-deployed robots collect data while they work, continually improving one shared policy together.
- Self-improvement自我提升A robot retrains itself on data from its own practice, getting steadily better with less reliance on human data.
10.12Inference-time gains and fast adaptation
After training ends, a model can still improve without changing weights much: more compute at inference, or fast adaptation to new tasks.
- Completing a new task by combining off-the-shelf models or algorithms, without updating any model parameters at all.
- Improving results by spending more compute at inference time, instead of changing the trained model.
- Best-of-N Sampling最优 N 采样Sampling N candidate outputs at once and using a scorer to pick the single highest-scoring one to actually execute.
- Value-Guided Sampling价值引导采样Sampling several candidate actions from a policy, then using a value function to score and pick the best one.
- Test-Time Training测试时训练Given new data at deployment, taking a few self-supervised update steps on it before predicting with the updated model.
- In-Context Learning上下文学习Learning a new task from a few examples given in the prompt, without updating the model's weights.
- Training a model on many tasks so it gets good at quickly learning new tasks, not just one fixed task.
- Training across many similar tasks so an agent can adapt to a new one with only a little trial and error.
- One-shot Imitation Learning单样本模仿学习The robot watches a new task demonstrated just once, then completes it starting from a different arrangement.