Foundations

Vanishing/Exploding Gradients

Vanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training.

Vanishing/exploding gradients are a failure mode of training deep neural networks. Since backpropagation multiplies local gradients along a chain of layers, if those local gradients are consistently below 1, their product shrinks toward zero across depth (vanishing) and early layers stop learning; if consistently above 1, the product explodes and training diverges. Fixes include better activation functions, weight initialization, batch/layer normalization, and residual connections.

How it works

The gradient reaching layer k of an N-layer network is a product of roughly N - k Jacobians. Products of numbers are unforgiving: with an average factor of 0.8, thirty layers deep leaves about 0.001 of the signal; with a factor of 1.2, it grows by more than 200x.

Two things set that factor. The activation derivative — sigmoid caps at 0.25, so it shrinks the signal at every layer by construction, while ReLU contributes exactly 1 on its positive side. And the weight scale, which is what He and Xavier initialization are designed to hold near the break-even point.

Residual connections sidestep the product entirely by adding an identity path whose derivative is 1.

When it breaks

  • Vanishing is silent. Training does not crash; the loss just plateaus at a mediocre value while early-layer weights barely move. Log per-layer gradient norms to see it.
  • Exploding is loud. The loss jumps to inf or NaN within a few steps. Gradient clipping by global norm (commonly 1.0) is the standard mitigation.
  • Sequence length, not just depth. An RNN backpropagates through every timestep, so a long sequence has the same effect as a very deep network — the original motivation for LSTMs and later for attention.
  • Mixed precision narrows the margin. float16 underflows below roughly 6e-8, so gradients that would merely be small in float32 become exactly zero without a loss scaler.

See also: Backpropagation, Activation Function

Learn more: Neural Networks & Backprop

Mentioned in

Lessons where this comes up in context.

On this page