Softmax
Softmax(归一化指数函数)CommonA function that turns a set of arbitrary real numbers into a probability distribution: all positive, summing to 1.
Softmax exponentiates each element of a real-valued vector, then divides by the sum of all the exponentials, producing a set of numbers that are all positive and sum to 1, so they can be treated as probabilities; larger input values get proportionally larger probabilities, and the exponential stretches the gaps between them. It has two main uses in deep learning. First, in a classification output layer, it turns the model's raw scores (logits) into per-class probabilities, which are then trained with cross-entropy loss. Second, inside attention mechanisms, it turns the similarity between a query and each key into attention weights. Dividing the scores by a temperature parameter before applying softmax lets you control how sharp the distribution is: a low temperature makes it close to picking only the largest value, while a high temperature makes it closer to uniform. RT-2 and OpenVLA, which discretize continuous actions into bins, also use softmax to give a probability to each bin.
ExampleA classifier scores 'cup,' 'bowl,' and 'plate' as 2.0, 1.0, and 0.1; softmax turns these into probabilities of roughly 0.66, 0.24, and 0.10.
- Also called
- Softmax Function, Temperature Softmax
- Related
- Cross-Entropy · Attention Mechanism · Self-Attention · Decoding Strategies · Spatial Softmax · Action Binning
- Sources
- 动手学深度学习:softmax 回归 (Chinese)
Dive into Deep Learning: Softmax Regression