Embodied AI Glossary中文

AdamW

AdamW 优化器Common

A version of the Adam optimizer that separates weight decay from the gradient update; the default choice for training most large models.

AdamW comes from Ilya Loshchilov and Frank Hutter's paper “Decoupled Weight Decay Regularization” (ICLR 2019). An optimizer is the algorithm that updates a model's parameters based on gradients, and Adam adapts the step size for each parameter individually. The older way to regularize Adam was to fold an L2 penalty into the gradient, but that penalty then gets rescaled by Adam's adaptive step sizes along with everything else, which weakens its effect. The authors showed that this L2-in-the-gradient approach is equivalent to weight decay (shrinking every parameter by a small proportion at each step) for plain SGD, but not for Adam — so they pulled the decay step out and applied it separately from the gradient update. This improves generalization and decouples the best decay value from the learning rate. AdamW is now one of the most common optimizers for training Transformer-based models.

ExampleIn PyTorch, torch.optim.AdamW(model.parameters(), lr=1e-4, weight_decay=0.01) uses AdamW directly; PyTorch's own defaults for it are a learning rate of 1e-3, betas of (0.9, 0.999), and a weight decay of 0.01.

Also called
Adam with Decoupled Weight Decay
Related
Optimizer · Gradient Descent · Learning Rate · Regularization · Learning Rate Schedule (Warmup and Cosine Decay)
Sources
Decoupled Weight Decay Regularization (arXiv 1711.05101)
PyTorch docs: torch.optim.AdamW

See it in the full glossary →