KL Regularization
KL 正则化AdvancedAdding a KL-divergence penalty to a training objective so the new policy doesn't drift too far from a reference.
KL regularization adds a term to the optimization objective — a coefficient times the KL divergence, a measure of how different two probability distributions are — that penalizes the current policy for drifting away from some reference policy. Three settings are common. In RLHF, InstructGPT adds a per-token KL penalty against the supervised fine-tuned model to stop the model from gaming the reward model. In offline reinforcement learning, methods such as BRAC use KL or similar divergences to keep the policy close to the behavior policy that generated the data, avoiding actions the data never covers. TRPO and PPO use KL to bound how much each update can change the policy. VLA reinforcement-learning fine-tuning commonly uses it too, to stop the policy from drifting and losing pretrained capabilities. Too large a coefficient stalls learning; too small and the constraint does nothing.
ExampleInstructGPT's PPO objective is “reward-model score minus β times log(π_RL / π_SFT)”: the larger β is, the less the new model dares to drift from the supervised fine-tuned model.
- Also called
- KL Penalty, KL Constraint
- Related
- Kullback-Leibler Divergence · Policy Constraint · Offline Reinforcement Learning · Reinforcement Learning from Human Feedback · Proximal Policy Optimization · Reward Hacking
- Sources
- Training language models to follow instructions with human feedback (InstructGPT, arXiv:2203.02155)
Behavior Regularized Offline Reinforcement Learning (BRAC, arXiv:1911.11361)