Activation Functions¶
What This Is¶
Activation functions are the element-wise nonlinearities a neural network inserts between linear layers. Without them a stack of linear layers collapses algebraically into a single linear layer — the activation is what gives depth its representational power.
The practical lesson is to start from the architecture's established activation—often ReLU-family functions in CNNs and GELU/SiLU-family functions in transformers—and match the output representation to the loss. The interesting part is the gradient-flow consequences and the diagnostic signs of failure.
When You Use It¶
- between every linear layer inside a neural network
- at the output, matched to the loss: linear + MSE, sigmoid + BCE, softmax + cross-entropy
- in gating structures: LSTM gates, attention softmax, gated feedforward
- at the end of a convolutional block before pooling
The Core Set¶
| Function | Formula | Range | Derivative | Notes |
|---|---|---|---|---|
| ReLU | max(0, x) |
[0, ∞) |
1 if x > 0 else 0 |
default for hidden layers |
| Leaky ReLU | max(αx, x), α ≈ 0.01 |
(-∞, ∞) |
1 or α |
reduces zero-gradient risk |
| ELU | x if x > 0 else α(e^x - 1) |
(-α, ∞) |
smooth | smoother alternative to leaky ReLU |
| GELU | x · Φ(x) |
(-∞, ∞) |
smooth | standard in transformers |
| Sigmoid | 1 / (1 + e^-x) |
(0, 1) |
σ(x)(1 - σ(x)), max 0.25 |
output activation for binary |
| Tanh | (e^x - e^-x) / (e^x + e^-x) |
(-1, 1) |
1 - tanh²(x), max 1 |
zero-centered, used in RNNs |
| Softmax | e^{x_i} / Σ_j e^{x_j} |
(0, 1) each, sum = 1 |
vector-valued Jacobian | output activation for multi-class |
| Swish / SiLU | x · σ(x) |
(-∞, ∞) |
smooth | common in modern vision backbones |
ReLU Is The Default — And Why¶
ReLU's local derivative is 1 on the positive side, so the activation itself does not shrink the gradient there. By contrast, each sigmoid derivative is at most 0.25; a path through 20 sigmoid derivatives contributes a factor no larger than 0.25^20 ≈ 10^-13 before weight matrices and other operations are considered.
For the underlying chain-rule picture, see Backpropagation.
Dead Neurons — The ReLU Failure Mode¶
If a ReLU unit's pre-activation is negative for every relevant example, its local derivative is zero and its incoming parameters receive no gradient through that unit. It can remain dead. A zero activation on one batch or example is normal ReLU sparsity, not proof of a dead unit.
Diagnostic: track units that remain zero across many representative batches and steps. There is no universal harmful percentage; compare against the architecture's expected activation sparsity and validation behavior.
Fixes:
- lower the learning rate — large updates often push units into the dead zone
- use Leaky ReLU, ELU, or GELU — these have nonzero gradient on the negative side
- use a proper initialization (see below) so pre-activations start balanced around zero
Vanishing Gradients — The Sigmoid / Tanh Failure Mode¶
Sigmoid saturates to 0 or 1 far from the origin; its derivative approaches zero in both tails. When inputs to a sigmoid are large in magnitude, the gradient through that unit is effectively zero and backpropagation cannot update upstream weights.
Tanh also saturates at ±1 but is zero-centered. Sigmoid and tanh remain useful in gates and recurrent state updates, while deep feed-forward stacks more often use non-saturating alternatives.
When you see sigmoid or tanh:
- output layer of a binary classifier (with BCE loss)
- RNN/LSTM gates, where saturation is actually desirable for a gate
- the output of a generator with a bounded range
- legacy code — check if ReLU would help
Output Activation Must Match The Loss¶
| Task | Output Activation | Loss |
|---|---|---|
| regression | none (linear) | MSE / MAE |
| binary classification | sigmoid probabilities for inference; logits during fused-loss training | binary cross-entropy |
| multi-class classification | softmax probabilities for inference; logits during fused-loss training | cross-entropy |
| multi-label classification | sigmoid probabilities per label; logits during fused-loss training | binary cross-entropy per label |
PyTorch note: nn.CrossEntropyLoss expects raw logits, not softmax outputs. Applying softmax before CrossEntropyLoss is a very common bug. Similarly nn.BCEWithLogitsLoss expects logits, not sigmoid outputs — and it is numerically safer than BCELoss(sigmoid(x), y).
Initialization Matches The Activation¶
Activation choice and weight initialization are coupled:
- ReLU-family → Kaiming (He) initialization: variance
2 / fan_in - tanh / sigmoid → Xavier (Glorot) initialization: variance
1 / fan_inor2 / (fan_in + fan_out)
PyTorch's nn.Linear uses a generic uniform initialization; it does not inspect the activation that follows. When initialization is a concern, initialize explicitly—for example, nn.init.kaiming_normal_(weight, nonlinearity="relu") for ReLU—and verify activation and gradient statistics. See the official nn.Linear documentation.
See Batch Normalization and Initialization for the full story.
Minimal Example¶
import torch.nn as nn
model = nn.Sequential(
nn.Linear(64, 128),
nn.ReLU(),
nn.Linear(128, 128),
nn.ReLU(),
nn.Linear(128, 10), # raw logits; no softmax here
)
loss_fn = nn.CrossEntropyLoss() # expects logits
No activation on the final layer — CrossEntropyLoss applies log_softmax internally.
What To Inspect¶
- the fraction of ReLU units that stay zero across representative batches and training steps
- gradient magnitudes per layer — shrinking by depth points at saturating nonlinearities
- pre-activation distributions — drift or saturation is evidence to inspect initialization, normalization, optimizer settings, and data scale
- whether the final layer matches the loss (softmax + CE both pre-applied is the classic double-softmax bug)
- what swapping ReLU for GELU or Swish does — usually small, sometimes meaningful
Failure Pattern¶
Applying softmax before CrossEntropyLoss. The loss expects logits and applies log_softmax internally; passing probabilities changes the objective and usually weakens the useful gradient signal.
A second failure pattern is placing sigmoid or tanh in hidden layers of a deep network and blaming the optimizer when gradients vanish. The activation is the cause.
A third failure pattern is ignoring ReLU units that remain inactive across the data. Persistent inactivity reduces usable capacity and can accompany an early plateau, though it is not the only possible cause.
A fourth failure pattern is ignoring the interaction between initialization and activation. A poor pairing can make activations or gradients shrink or grow with depth; inspect them instead of assuming failure after a fixed number of batches.
Quick Checks¶
- Is the final layer activation matched to the loss?
- Do the hidden activations match the architecture and show healthy activation/gradient statistics?
- Is the initialization matched to the activation family?
- Which ReLU units remain zero across many representative batches and steps?
- Are gradient magnitudes roughly comparable across layers?
Practice¶
- Train the same MLP with sigmoid, tanh, ReLU, and GELU hidden activations. Plot loss curves and compare.
- Stack 20 layers with sigmoid and ReLU respectively and plot gradient magnitudes per layer.
- Count dead ReLUs after a few epochs with a too-large learning rate. Repeat with a smaller learning rate.
- Swap ReLU for Leaky ReLU in the dead-neuron scenario and observe the recovery.
- Deliberately apply softmax before
CrossEntropyLoss; compare gradients, convergence, and validation performance with the correct logits input. - Explain why GELU is common in transformers and why ReLU is still common in CNNs.
- Describe one reason tanh survives in LSTM cells while being unusual in feedforward nets.
- Explain why the output activation for multi-label classification is sigmoid, not softmax.
- State the relationship between activation choice and initialization scheme.
- Describe what happens to gradients when a sigmoid is saturated at 0.99.
Runnable Example¶
Run the small MLP recipe from the repository root:
.venv/bin/python examples/deep-learning-recipes/mlp_training_recipe.py
Use the activation substitutions on this page as controlled sketches: change only the hidden activation, then compare convergence and the fraction of zero activations.
Longer Connection¶
Activation functions sit next to:
- Backpropagation — the chain-rule picture that makes dead neurons and vanishing gradients concrete
- Batch Normalization and Initialization — the layers and init schemes that keep pre-activations in a healthy range
- Optimizers and Regularization — the update rules that consume the gradient through the activation
- Debugging Deep Learning — how to locate the broken layer when training misbehaves
The activation is the nonlinear knob that makes depth work. Picking it is easy; diagnosing when it misbehaves is where the skill lives.