GELU & Swish Activation Calculator

GELU, Swish & Mish

Modern smooth activation functions used in Transformers

GELU
x · Φ(x)
Used in BERT, GPT
Swish / SiLU
x · σ(x)
Used in EfficientNet
Mish
x · tanh(softplus(x))
Used in YOLOv4
0.000000
0.000000
0.000000

Comparison

Formulas

GELU
x * CDF(x)
Approx: 0.5x(1 + tanh(sqrt(2/pi)(x + 0.044715x^3)))
Swish / SiLU
x * sigmoid(x)
= x / (1 + e^(-x))
Mish
x * tanh(softplus(x))
= x * tanh(ln(1 + e^x))

Key Properties

  • • All are smooth and differentiable
  • • Non-monotonic (allow negative values)
  • • Self-gating mechanism
  • • No dying neuron problem

When to Use

  • GELU: Transformers, NLP models
  • Swish: Image classification, EfficientNet
  • Mish: Object detection, YOLO

What is the GELU activation function?

GELU (Gaussian Error Linear Unit) is an activation function defined as GELU(x) = x × Φ(x), where Φ(x) is the cumulative distribution function of the standard normal distribution. Unlike ReLU which has a hard cutoff at zero, GELU applies a smooth, probabilistic gating mechanism that weights inputs by their magnitude. It can be approximated as GELU(x) ≈ 0.5x(1 + tanh(√(2/π)(x + 0.044715x³))). GELU is the default activation in transformer architectures including BERT, GPT, and Vision Transformers (ViT). The closely related Swish function, defined as Swish(x) = x × σ(x), was discovered independently by Google Brain and shares similar properties. Both outperform ReLU in transformer-based models by allowing small negative values to pass through, improving gradient flow in deep architectures.

Built with care by Alpiaal