Softmax Function Calculator

Softmax Function

softmax(xi) = exi / Σexj
1.00
Sharp (0.1)Normal (1.0)Smooth (5.0)

Input Logits

Output Probabilities

P(x1)
65.90%
P(x2)
24.24%
P(x3)
9.86%
Sum: 1.000000 (should be 1.0)

Probability Distribution

Formulas

Standard Softmax
P(x_i) = e^(x_i) / sum(e^(x_j))
With Temperature
P(x_i) = e^(x_i/T) / sum(e^(x_j/T))
Numerically stable: subtract max(x) before exp to prevent overflow

Temperature Effect

  • τ → 0: One-hot output (argmax)
  • τ = 1: Standard softmax
  • τ → ∞: Uniform distribution

Use Cases

  • • Multi-class classification output
  • • Attention mechanisms
  • • Reinforcement learning policies
  • • Knowledge distillation (with temp)

What is the softmax function and how do you calculate it?

The softmax function converts a vector of real numbers (logits) into a probability distribution. For a vector z of K elements, softmax(z_i) = e^(z_i) / Σ(e^(z_j)) for j = 1 to K. Each output value is between 0 and 1, and all outputs sum to exactly 1, making softmax the standard choice for the output layer of multi-class classification neural networks. Temperature scaling divides the logits by a temperature parameter T before applying softmax: a lower temperature (T < 1) makes the distribution sharper (more confident), while a higher temperature (T > 1) produces a softer, more uniform distribution. Softmax is used in attention mechanisms in transformers, reinforcement learning policy networks, and any model that needs to output class probabilities across mutually exclusive categories.

Built with care by Alpiaal