Regularization
正则化CommonAdding constraints or penalties during training to keep a model from memorizing training data, improving performance on new data.
Regularization is a broad category of techniques for preventing overfitting: whenever a model keeps improving on the training set but gets worse on new data, regularization is used to limit the model's effective complexity. Explicit regularization adds a penalty term to the loss function, such as L2 regularization (also called weight decay, which penalizes the sum of squared weights and shrinks them overall) and L1 regularization (which penalizes absolute value and pushes many weights to exactly zero); implicit regularization includes early stopping, dropout (randomly zeroing some neurons during training), and data augmentation. Reinforcement learning has several regularizers of its own: entropy regularization encourages the policy to stay somewhat random to keep exploring, KL regularization keeps a fine-tuned policy from drifting too far from the original model, and behavior regularization keeps an offline reinforcement-learning policy close to the actions in the dataset.
ExampleTraining a visuomotor policy applies random cropping to camera images (data augmentation) and sets weight decay in the AdamW optimizer; RLHF subtracts a KL penalty term from the reward, keeping the model from drifting away from the original language model just to rack up a higher score.
- Also called
- Regularization Term
- Related
- Overfitting · Dropout · Early Stopping · Data Augmentation · Entropy Regularization · KL Regularization
- Sources
- Wikipedia: Regularization (mathematics)
Dive into Deep Learning: Weight Decay