ReLU & Leaky ReLU Calculator

ReLU Variants

max(0, x)
0.0000
0.0000

Visualization

Formulas

ReLU
f(x) = max(0, x)
f'(x) = 1 if x > 0, else 0
Leaky ReLU
f(x) = max(ax, x)
f'(x) = 1 if x > 0, else a
ELU
f(x) = x if x >= 0, a(e^x - 1) otherwise
f'(x) = 1 if x > 0, else a*e^x

Comparison

VariantDying ReLUSmooth
ReLUYesNo
Leaky ReLUNoNo
ELUNoYes

When to Use

  • ReLU: Default choice, fast computation
  • Leaky ReLU: When experiencing dead neurons
  • ELU: Better convergence, slightly slower

What is the ReLU activation function?

ReLU (Rectified Linear Unit) is the most widely used activation function in deep learning, defined as f(x) = max(0, x). It outputs the input directly if positive and zero otherwise. ReLU solved the vanishing gradient problem that plagued sigmoid and tanh activations, enabling training of much deeper networks. Leaky ReLU is a variant defined as f(x) = max(αx, x) where α is a small constant (typically 0.01), which prevents "dying ReLU" neurons by allowing a small gradient for negative inputs. Other variants include Parametric ReLU (PReLU) where α is learned during training, and ELU which uses an exponential for negative values. ReLU is computationally efficient and is the default choice for hidden layers in convolutional and fully connected neural networks.

Built with care by Alpiaal