Backpropagation
Backpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network.
Backpropagation ("backprop") is the algorithm that makes training deep neural networksNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. computationally feasible. It applies the chain rule from calculus, working backward from the loss through the network's computation graph, to compute the exact gradient of the loss with respect to every parameter in roughly the same amount of work as one forward pass — regardless of how many parameters the network has. Modern frameworks like PyTorchPyTorchPyTorch is the dominant deep learning framework in both research and production, providing tensor computation, GPU dispatch, and automatic differentiation. implement this automatically as autograd.
How it works
The forward pass records every operation into a directed graph and keeps
the intermediate activations. The backward pass then walks that graph in
reverse, starting from dL/dL = 1, and at each node multiplies the
incoming gradient by that operation's local derivative — reverse-mode
automatic differentiation.
For a linear layer y = Wx + b with incoming gradient g = dL/dy:
dL/dW = g xᵀdL/db = gdL/dx = Wᵀ g, which is passed further back
Because dL/dx is computed once per node and reused by all its inputs,
the whole gradient costs roughly one extra forward pass, not one pass per
parameter. The resulting gradients are then consumed by
gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient..
When it breaks
- Memory, not compute, is the limit. Every stored activation is held until its backward pass runs, so memory scales with depth times batch size. Gradient checkpointing trades recomputation for memory.
- Long chains degrade. Multiplying many local derivatives produces vanishing or exploding gradientsVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training.; residual connectionsResidual ConnectionA residual (skip) connection adds a layer's input back to its output, giving gradients a direct path backward and making very deep networks trainable. and normalization exist mainly to counter this.
- Silent graph breaks. Detaching a tensor, converting to NumPy, or doing an in-place edit severs the graph, and the affected parameters simply stop updating with no error raised.
- Stale gradients. Gradients accumulate by default; forgetting
optimizer.zero_grad()sums this step's gradient onto the last one.
See also: Gradient DescentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient., Neural NetworkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation.
Learn more: Neural Networks & Backprop · Wikipedia: Backpropagation
Mentioned in
Lessons where this comes up in context.
- Applied & Agentic SystemsHow prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- ML FundamentalsSupervised/unsupervised learning, loss functions, gradient descent
- Neural Networks & BackpropFrom Karpathy's micrograd approach — building a tiny neural net and stepping through forward/backward passes
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
Gradient Descent
Gradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient.
Vanishing/Exploding Gradients
Vanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training.