The softmax function converts a vector of real numbers (logits) into a probability distribution. For a vector z of K elements, softmax(z_i) = e^(z_i) / Σ(e^(z_j)) for j = 1 to K. Each output value is between 0 and 1, and all outputs sum to exactly 1, making softmax the standard choice for the output layer of multi-class classification neural networks. Temperature scaling divides the logits by a temperature parameter T before applying softmax: a lower temperature (T < 1) makes the distribution sharper (more confident), while a higher temperature (T > 1) produces a softer, more uniform distribution. Softmax is used in attention mechanisms in transformers, reinforcement learning policy networks, and any model that needs to output class probabilities across mutually exclusive categories.
Built with care by Alpiaal