| Property | Tanh | Sigmoid |
|---|---|---|
| Range | (-1, 1) | (0, 1) |
| Zero-centered | Yes | No |
| Max derivative | 1.0 | 0.25 |
The hyperbolic tangent (tanh) function is defined as tanh(x) = (e^x − e^(−x)) / (e^x + e^(−x)), mapping inputs to the range (−1, 1). It is a rescaled version of the sigmoid function: tanh(x) = 2σ(2x) − 1. Tanh is preferred over sigmoid in hidden layers because it is zero-centered, meaning its outputs have a mean closer to zero, which helps gradients flow more effectively during backpropagation. The derivative is tanh'(x) = 1 − tanh²(x), with a maximum of 1.0 at x = 0. Tanh is commonly used in recurrent neural networks (RNNs), LSTM cell state updates, and as the activation in hidden layers of shallow networks. Like sigmoid, it suffers from vanishing gradients for large input magnitudes, which is why ReLU has replaced it in most deep architectures.
Built with care by Alpiaal