Modern smooth activation functions used in Transformers
GELU (Gaussian Error Linear Unit) is an activation function defined as GELU(x) = x × Φ(x), where Φ(x) is the cumulative distribution function of the standard normal distribution. Unlike ReLU which has a hard cutoff at zero, GELU applies a smooth, probabilistic gating mechanism that weights inputs by their magnitude. It can be approximated as GELU(x) ≈ 0.5x(1 + tanh(√(2/π)(x + 0.044715x³))). GELU is the default activation in transformer architectures including BERT, GPT, and Vision Transformers (ViT). The closely related Swish function, defined as Swish(x) = x × σ(x), was discovered independently by Google Brain and shares similar properties. Both outperform ReLU in transformer-based models by allowing small negative values to pass through, improving gradient flow in deep architectures.
Built with care by Alpiaal