Foundations

Backpropagation

Backpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network.

Backpropagation ("backprop") is the algorithm that makes training deep neural networks computationally feasible. It applies the chain rule from calculus, working backward from the loss through the network's computation graph, to compute the exact gradient of the loss with respect to every parameter in roughly the same amount of work as one forward pass — regardless of how many parameters the network has. Modern frameworks like PyTorch implement this automatically as autograd.

How it works

The forward pass records every operation into a directed graph and keeps the intermediate activations. The backward pass then walks that graph in reverse, starting from dL/dL = 1, and at each node multiplies the incoming gradient by that operation's local derivative — reverse-mode automatic differentiation.

For a linear layer y = Wx + b with incoming gradient g = dL/dy:

  • dL/dW = g xᵀ
  • dL/db = g
  • dL/dx = Wᵀ g, which is passed further back

Because dL/dx is computed once per node and reused by all its inputs, the whole gradient costs roughly one extra forward pass, not one pass per parameter. The resulting gradients are then consumed by gradient descent.

When it breaks

  • Memory, not compute, is the limit. Every stored activation is held until its backward pass runs, so memory scales with depth times batch size. Gradient checkpointing trades recomputation for memory.
  • Long chains degrade. Multiplying many local derivatives produces vanishing or exploding gradients; residual connections and normalization exist mainly to counter this.
  • Silent graph breaks. Detaching a tensor, converting to NumPy, or doing an in-place edit severs the graph, and the affected parameters simply stop updating with no error raised.
  • Stale gradients. Gradients accumulate by default; forgetting optimizer.zero_grad() sums this step's gradient onto the last one.

See also: Gradient Descent, Neural Network

Learn more: Neural Networks & Backprop · Wikipedia: Backpropagation

Mentioned in

Lessons where this comes up in context.

On this page