Activation Function
An activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function.
An activation function is the nonlinear function applied after a
neural networkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. neuron's weighted sum —
without it, stacking layers would collapse into one large linear function,
no more expressive than a single layer. ReLU (max(0, z)) is the
standard modern default: cheap, and its derivative doesn't shrink for
positive inputs, avoiding much of the vanishing gradientVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training.
problem that plagued older choices like sigmoid and tanh. GELU,
a smoother ReLU variant, is used in most modern transformers.
How it works
A neuron computes z = w · x + b, then emits a = f(z). The choice of
f decides what gradient flows back through that unit during
backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network., because the chain rule
multiplies by f'(z) at every layer.
- Sigmoid squashes to
(0, 1); its derivative peaks at0.25and decays toward zero for large|z|. - Tanh squashes to
(-1, 1), zero-centered, derivative peaks at1. - ReLU passes
zthrough unchanged when positive, sof'(z) = 1there and0otherwise — no shrinkage, and a cheap comparison instead of an exponential. - GELU and SiLU/Swish are smooth approximations of ReLU with
nonzero gradient for small negative
z.
Without any nonlinearity, composing W₂(W₁x) is just (W₂W₁)x, a
single linear map.
When it breaks
- Dead ReLUs. A unit whose pre-activation is negative for every
input has zero gradient forever and never recovers. Usually caused by
a large learning rate or a badly scaled
initializationWeight InitializationWeight initialization sets a neural network's starting parameter values, using schemes like Xavier/He initialization to keep gradients well-scaled from the first training step..
LeakyReLUorGELUavoids this by keeping a small negative-side slope. - Saturation. Sigmoid and tanh flatten for large
|z|, so deep stacks of them stall — the historical reason sigmoid networks were hard to train past a few layers. - Mismatched output activation. Applying softmax or sigmoid yourself
and also passing logits to a loss that applies it internally (e.g.
CrossEntropyLossin PyTorch) double-applies the nonlinearity and quietly degrades training. - Non-zero-centered outputs. ReLU emits only non-negative values, which biases gradient directions; normalization layers usually absorb this.
See also: Neural NetworkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation., Vanishing GradientsVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training.
Learn more: Neural Networks & Backprop
Mentioned in
Lessons where this comes up in context.
Neural Network
A neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation.
Gradient Descent
Gradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient.