| Variant | Dying ReLU | Smooth |
|---|---|---|
| ReLU | Yes | No |
| Leaky ReLU | No | No |
| ELU | No | Yes |
ReLU (Rectified Linear Unit) is the most widely used activation function in deep learning, defined as f(x) = max(0, x). It outputs the input directly if positive and zero otherwise. ReLU solved the vanishing gradient problem that plagued sigmoid and tanh activations, enabling training of much deeper networks. Leaky ReLU is a variant defined as f(x) = max(αx, x) where α is a small constant (typically 0.01), which prevents "dying ReLU" neurons by allowing a small gradient for negative inputs. Other variants include Parametric ReLU (PReLU) where α is learned during training, and ELU which uses an exponential for negative values. ReLU is computationally efficient and is the default choice for hidden layers in convolutional and fully connected neural networks.
Built with care by Alpiaal