Embodied AI Glossary中文

Vanishing / Exploding Gradients

梯度消失 / 梯度爆炸Common

Gradients shrinking to near zero or blowing up as they're multiplied across many layers during backpropagation, stalling deep-network training.

A deep network computes gradients through backpropagation, and the gradient reaching an early layer is a product of the derivatives (matrices) of every layer after it. When those factors tend to be small, the gradient shrinks toward zero by the time it reaches shallow layers, so the early layers barely learn at all — vanishing gradients; when the factors tend to be large, the product grows exponentially, and a single update can send parameters diverging and the loss to NaN — exploding gradients. Bengio and colleagues analyzed this problem in recurrent neural networks in 1994, and Pascanu and colleagues proposed gradient-norm clipping to handle explosion in 2013. Other common fixes include activation functions that don't saturate easily, like ReLU (sigmoid's derivative approaches zero when its input is very large or very small), Xavier/He initialization, residual connections, normalization layers, and LSTM's gating structure. The gradient clipping commonly set when training VLAs and diffusion policies is also there to prevent this kind of numerical instability.

ExampleStacking dozens of fully connected layers with sigmoid activations, the gradient in the first few layers is often close to 0 and their parameters barely move; switching to ReLU and adding residual connections lets a network of the same depth train normally.

Also called
Vanishing Gradient, Exploding Gradient
Related
Backpropagation · Gradient Clipping · Residual Network · Activation Function · Normalization Layers · Long Short-Term Memory / Gated Recurrent Unit
Sources
Dive into Deep Learning: Numerical Stability and Initialization
On the difficulty of training Recurrent Neural Networks (Pascanu et al., arXiv 1211.5063)

See it in the full glossary →