Exponential Moving Average
指数移动平均EMAAdvancedA weighted average of past values that decays exponentially, often used to get smoother, more stable model weights.
EMA is a weighted average updated at every step as “average ← β × average + (1 − β) × current value,” where β is the decay rate — the closer to 1, the longer the effective averaging window. The most common use in deep learning is keeping an EMA copy of the model's weights: training keeps updating the original weights as usual, while evaluation and deployment switch to the smoother EMA weights, which are usually more stable; diffusion models depend on this especially heavily, with the DDPM paper using a decay rate of 0.9999. In reinforcement learning, the soft update of a target network (Polyak averaging) is also an EMA, keeping the Q-value target changing more smoothly. The Adam optimizer's running estimates of a gradient's first and second moments are EMAs too.
ExampleDiffusion Policy's official code enables EMA by default, with the decay rate ramping up from 0 to a cap of 0.9999 over training, and evaluation rollouts use the EMA model rather than the raw weights.
- Also called
- EMA, Polyak Averaging
- Related
- Target Network · Diffusion Policy · Checkpoint · Optimizer · Denoising Diffusion Probabilistic Model · Soft Actor-Critic
- Sources
- Denoising Diffusion Probabilistic Models (arXiv:2006.11239)
diffusion_policy: train_diffusion_unet_image_workspace.yaml
OpenAI Spinning Up: Soft Actor-Critic