Vanishing/Exploding Gradients
Vanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training.
Vanishing/exploding gradients are a failure mode of training deep neural networksNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation.. Since backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network. multiplies local gradients along a chain of layers, if those local gradients are consistently below 1, their product shrinks toward zero across depth (vanishing) and early layers stop learning; if consistently above 1, the product explodes and training diverges. Fixes include better activation functionsActivation FunctionAn activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function., weight initializationWeight InitializationWeight initialization sets a neural network's starting parameter values, using schemes like Xavier/He initialization to keep gradients well-scaled from the first training step., batch/layer normalizationBatch/Layer NormalizationBatch and layer normalization rescale a layer's outputs during training to keep values well-behaved as they pass through many stacked layers., and residual connections.
How it works
The gradient reaching layer k of an N-layer network is a product of
roughly N - k Jacobians. Products of numbers are unforgiving: with an
average factor of 0.8, thirty layers deep leaves about 0.001 of the
signal; with a factor of 1.2, it grows by more than 200x.
Two things set that factor. The activationActivation FunctionAn activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function.
derivative — sigmoid caps at 0.25, so it shrinks the signal at every
layer by construction, while ReLU contributes exactly 1 on its positive
side. And the weight scale, which is what
He and Xavier initializationWeight InitializationWeight initialization sets a neural network's starting parameter values, using schemes like Xavier/He initialization to keep gradients well-scaled from the first training step. are
designed to hold near the break-even point.
Residual connectionsResidual ConnectionA residual (skip) connection adds a layer's input back to its output, giving gradients a direct path backward and making very deep networks trainable. sidestep the
product entirely by adding an identity path whose derivative is 1.
When it breaks
- Vanishing is silent. Training does not crash; the loss just plateaus at a mediocre value while early-layer weights barely move. Log per-layer gradient norms to see it.
- Exploding is loud. The loss jumps to
inforNaNwithin a few steps. Gradient clipping by global norm (commonly1.0) is the standard mitigation. - Sequence length, not just depth. An RNNRNN (Recurrent Neural Network)An RNN is a neural network for sequences that processes input one step at a time, carrying a hidden state forward — the main predecessor to transformers. backpropagates through every timestep, so a long sequence has the same effect as a very deep network — the original motivation for LSTMs and later for attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer..
- Mixed precision narrows the margin.
float16underflows below roughly6e-8, so gradients that would merely be small infloat32become exactly zero without a loss scaler.
See also: BackpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network., Activation FunctionActivation FunctionAn activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function.
Learn more: Neural Networks & Backprop
Mentioned in
Lessons where this comes up in context.
Backpropagation
Backpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network.
Weight Initialization
Weight initialization sets a neural network's starting parameter values, using schemes like Xavier/He initialization to keep gradients well-scaled from the first training step.