Normalization Layers
归一化层(层归一化 / RMSNorm / 批归一化)LN / RMSNorm / BNAdvancedLayers that rescale a network's intermediate features back into a stable numeric range, making deep networks train faster and more stably.
A normalization layer rescales intermediate features by something like 'subtract the mean, divide by the standard deviation,' then multiplies by a learnable scale and adds a learnable shift, keeping values from spiraling out of control as depth increases and making training more stable. Batch normalization (BN, 2015) computes each channel's mean and variance across a batch, and became standard in convolutional networks, but it depends on batch size and behaves differently during training versus inference. Layer normalization (LN, 2016) instead computes statistics across all of a single sample's features, behaving the same way in training and inference, which made it the standard for Transformers. RMSNorm (2019) drops the mean-subtraction step and only divides by the root-mean-square, saving compute; large models like Llama use it, so VLAs built on them as a language backbone mostly use RMSNorm too. Diffusion Transformers also commonly use adaptive layer normalization to inject conditions like the timestep.
ExampleDiffusion Policy replaces every batch normalization layer in its ResNet-18 vision encoder with group normalization (GroupNorm), because batch normalization was found to train unstably when combined with the exponential moving average (EMA) weights diffusion models commonly use.
- Also called
- LayerNorm, RMSNorm, BatchNorm
- Related
- Adaptive Layer Normalization · Transformer · Residual Network · Convolutional Neural Network · Llama · Exponential Moving Average
- Sources
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift (arXiv:1502.03167)
Layer Normalization (arXiv:1607.06450)
Root Mean Square Layer Normalization (arXiv:1910.07467)