Activation Function
激活函数CommonThe nonlinear function after each layer's linear transform, letting a network fit complex relationships.
An activation function follows every linear transform in a neural network's layers; without it, stacking layers would still amount to just one linear transform, no matter how many layers there are. A few are common: ReLU zeroes out negative numbers and passes positive ones through unchanged, is cheap to compute, and was the default choice through the convolutional-network era. GELU, defined by Hendrycks and Gimpel in 2016 as x times Φ(x), where Φ is the standard normal distribution's cumulative distribution function, is common in Transformers such as BERT and ViT. SiLU, also called Swish, is x times sigmoid(x). SwiGLU, proposed by Shazeer in 2020, is a gated-linear-unit variant that multiplies two linear projections element-wise, with one of them passed through Swish first. Llama, for instance, replaced ReLU with SwiGLU in its Transformer feed-forward layers.
ExampleFor inputs −1, 0.5, and 2, ReLU outputs 0, 0.5, and 2; GELU outputs about −0.16, 0.35, and 1.95, no longer chopping negative numbers straight down to zero.
- Also called
- ReLU, GELU, SiLU, SwiGLU
- Related
- Multilayer Perceptron · Transformer · Neural Network · Normalization Layers · Vanishing / Exploding Gradients · Llama
- Sources
- Gaussian Error Linear Units (GELUs) (arXiv:1606.08415)
GLU Variants Improve Transformer (arXiv:2002.05202)
LLaMA: Open and Efficient Foundation Language Models (arXiv:2302.13971)