Neural Networks & Backprop
From Karpathy's micrograd approach — building a tiny neural net and stepping through forward/backward passes
ML Fundamentals established the loop: compute a loss, compute its gradient with respect to the parameters, step against the gradient. This lesson is about the one piece left unexplained — how you actually compute that gradient when your model is a neural networkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. with millions of parameters chained through dozens of layers. The answer is an algorithm called backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network., and understanding it from scratch (the micrograd approach) demystifies everything that follows.
What a neuron actually is
A single artificial neuron is almost embarrassingly simple: it takes some inputs, multiplies each by a weight, adds them up with a bias term, and passes the result through a nonlinear function:
z = w1*x1 + w2*x2 + ... + wn*xn + b
a = activation(z)The weights (w1...wn) and bias (b) are the learnable parameters.
The activation functionActivation FunctionAn activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function. (commonly ReLU(z) = max(0, z), or historically
tanh / sigmoid) introduces nonlinearity — without it, stacking layers of
neurons would collapse into one big linear function, no more expressive
than a single layer. Nonlinearity is what lets networks represent curved
decision boundaries and complex functions at all.
A layer is a bunch of these neurons run in parallel on the same input. A network is layers stacked so each layer's output feeds the next layer's input. That's the entire architecture of a basic feedforward neural network, also called a multi-layer perceptronPerceptronThe Perceptron (1958) was the first learning system built from an artificial neuron, and the direct ancestor of the neural network training loop. (MLP) — the complexity in modern AIAI (Artificial Intelligence)AI is the field of building systems that perform tasks normally requiring human intelligence — reasoning, perception, language, and decision-making. comes from what's connected to what and how (see Computer Vision and AttentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. & TransformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.), not from the neuron itself.
Why any of this can represent complex functions
A natural question: why does chaining together these simple weighted-sum-plus-nonlinearity units let you represent essentially anything? The universal approximation theoremUniversal Approximation TheoremThe universal approximation theorem proves that a feedforward network with even one hidden layer can approximate any continuous function, given enough neurons. answers this formally: a feedforward network with even a single hidden layer, given enough neurons in that layer, can approximate any continuous function to arbitrary precision. This is a real, proven result — not marketing language.
The catch is that "enough neurons" can be an impractically large number for a single layer, and the theorem says nothing about whether such a network is learnable via gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. in practice. This is exactly why depth matters more than the theorem alone would suggest: a deeper network can represent the same function with far fewer total parameters, and turns out to be far more trainable in practice — see "Where the 'deep' in deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones. comes from," below.
Choosing an activation function
Not all nonlinearities are equal in practice, and the choice has a real history behind it:
- Sigmoid (
1 / (1 + e^-z)) and tanh were the original defaults — smooth, bounded outputs. Their flaw: for large positive or negativez, the function flattens out almost completely, so its derivative is close to zero. Chained across many layers (see vanishing gradients, below), this made deep networks with these activations extremely slow or impossible to train. - ReLU (
max(0, z)) fixed this for positive inputs — its derivative is exactly 1 whereverz > 0, so it doesn't shrink gradients flowing through active neurons. It's cheap to compute and was a major factor in making deep networks trainable at all. Its own flaw: a neuron whose input is always negative outputs zero forever and stops learning (the "dying ReLU" problem). - GELU and other smooth ReLU variants (used in most modern transformers, including the ones in Attention & Transformers) soften ReLU's sharp corner at zero, which empirically trains slightly better at the scale modern models operate at, at a small extra compute cost.
| Activation | Derivative behavior | Main failure mode |
|---|---|---|
| Sigmoid | Near zero for large inputs | Vanishing gradients with depth |
| Tanh | Near zero for large inputs | Vanishing gradients with depth |
| ReLU | Exactly 1 when z > 0, else 0 | Dying neurons stuck at zero |
| GELU | Smooth, nonzero near zero | Slightly more compute per unit |
The forward pass
Running an input through the network to produce an output is called the
forward pass. Concretely, for a tiny 2-layer network computing a
prediction y_pred from an input x:
h = ReLU(W1 · x + b1) # hidden layer
y_pred = W2 · h + b2 # output layer
loss = (y_pred - y_true)^2 # e.g. squared error against the true labelEvery one of those operations — multiply, add, ReLU, subtract, square — is a simple, differentiable function. That fact is the entire key to what comes next.
The same graph is traversed twice: values flow left to right to produce a loss, then gradients flow right to left back to the parameters.
The chain rule, and why it scales
To do gradient descent we need ∂loss/∂W1, ∂loss/∂b1, ∂loss/∂W2,
∂loss/∂b2 — how the loss changes with respect to every parameter,
even ones buried several operations deep. Computing each of these from
scratch, independently, would be enormously wasteful — many of them share
intermediate work.
Backpropagation is just the chain rule from calculus, applied
systematically: if loss depends on y_pred which depends on h which
depends on W1, then
∂loss/∂W1 = (∂loss/∂y_pred) · (∂y_pred/∂h) · (∂h/∂W1)The trick is order of computation: work backward from the loss, computing each local derivative once and reusing it for everything upstream. That's why it's called backpropagation — you build a computation graph forward, then walk it backward, accumulating gradients.
Write the two-layer network above with explicit pre-activations. With , , , and a scalar loss , define the error signal at each layer as the gradient of the loss with respect to that layer's pre-activation:
where is elementwise multiplication. Once you have for a layer, that layer's parameter gradients are immediate:
with the incoming activation (). This is the whole algorithm: compute at the output, push it backward with and the local derivative , and read off parameter gradients as an outer product at each stop. The transpose appears because the forward pass maps activations through , so the derivative maps sensitivities back through .
Note what depends on: the activation from the forward pass. Those values must still be in memory when the backward pass reaches that layer, which is the cost discussed below.
Stepping through it, node by node
This is the exercise worth actually doing once (as Karpathy's micrograd does): build a tiny scalar "autograd" engine where every number remembers (a) the operation that produced it and (b) its inputs. Two operations, concretely:
# forward: c = a * b
# backward, given dL/dc (gradient flowing in from later in the graph):
# dL/da = dL/dc * b
# dL/db = dL/dc * a
# forward: c = a + b
# backward, given dL/dc:
# dL/da = dL/dc * 1
# dL/db = dL/dc * 1Every operation you'll ever use (multiply, add, ReLU, exp, division, matrix multiply for real networks) has a similarly simple local gradient rule. Backprop is just: run the forward pass while recording the graph, then walk it in reverse, multiplying local gradients along each path and summing where paths merge (a value used twice contributes to the gradient twice). At the end, every parameter in the network has an exact gradient of the loss with respect to it — computed in roughly the same amount of work as the forward pass itself, regardless of how many parameters there are.
That last property — the gradient of all parameters costs about the same as one forward pass — is what makes training networks with billions of parameters computationally feasible at all. Without it, gradient descent at LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. scale would be a non-starter.
You could estimate gradients without any calculus: nudge one parameter by a tiny amount, re-run the network, and see how much the loss changed. That's a finite difference, and it needs one forward pass per parameter.
For a small MLP with 1M parameters:
- Finite differences: 1,000,000 forward passes per gradient step
- Backprop: 1 forward + 1 backward ≈ 3 forward passes' worth of work
That's a factor of roughly 300,000 — and the gap grows linearly with parameter count, so at a billion parameters it is a factor of hundreds of millions. Backprop is not a speedup over the naive method; it is the difference between possible and impossible.
The price is memory. Take a 10-layer MLP with 1,024 units per layer, batch size 64, stored as 4-byte floats: each layer's activations are 64 × 1024 × 4 ≈ 0.26 MB, so all ten together are only ~2.6 MB. Tiny here — but the same accounting at transformer scale is why training a model needs far more memory than running it.
A fully worked numeric example
Concrete numbers make this stick better than symbols alone. Take the
smallest possible case: loss = (a * b + c), with a = 2, b = -3,
c = 10.
Forward pass:
d = a * b = 2 * -3 = -6
loss = d + c = -6 + 10 = 4Backward pass, starting from dL/dL = 1 (the loss's gradient with
respect to itself) and working backward using the local rules above:
dL/dd = dL/dloss * 1 = 1 # loss = d + c, local grad w.r.t. d is 1
dL/dc = dL/dloss * 1 = 1 # local grad w.r.t. c is 1
dL/da = dL/dd * b = 1 * -3 = -3 # d = a*b, local grad w.r.t. a is b
dL/db = dL/dd * a = 1 * 2 = 2 # local grad w.r.t. b is aRead dL/da = -3 as: "increasing a by a small amount increases the loss
by about 3 times that amount" — so gradient descent will decrease a.
Step through this exact example node by node, forward then backward —
try changing a, b, or c and re-running it:
Step 0 / 5: Start
Set a, b, c, then step through the forward pass, then backward.
Every gradient in a million-parameter network is computed by chaining
exactly these two rules (multiply, add) — plus a handful of others for
whatever operations appear in the graph — over and over, automatically.
This is precisely what loss.backward() does in a real framework (see
Tooling & The Dev Stack) — the manual
version above is what autograd is doing under the hood.
Where the "deep" in deep learning comes from
"Depth" just means stacking many layers. Deeper networks can represent more complex functions with fewer total parameters than a single very wide layer — each layer builds on features the previous layer extracted, rather than trying to learn the whole mapping at once. Practically, depth introduces its own problems, and three specific fixes are worth knowing by name, since all three show up throughout this course:
- Vanishing/exploding gradientsVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training.: backprop multiplies local gradients along a chain (as in the worked example above). If those local gradients are consistently a bit less than 1, their product shrinks toward zero across many layers — early layers get almost no gradient signal and stop learning. If they're consistently a bit more than 1, the product explodes instead, and training diverges. This is the practical reason sigmoid/tanh fell out of favor for deep networks (see activation functions, above) — their small derivatives compound badly across depth.
- Weight initializationWeight InitializationWeight initialization sets a neural network's starting parameter values, using schemes like Xavier/He initialization to keep gradients well-scaled from the first training step.: the starting values of the weights matter more than intuition suggests — initializing too large or too small makes vanishing/exploding gradients worse from step one. Schemes like Xavier/Glorot initialization and He initialization set the initial random weight scale based on the number of inputs/outputs of each layer, specifically to keep activations and gradients at a consistent scale as they pass through the network.
- Normalization layers (batch normalization, layer normalization): rescale a layer's outputs (mean roughly 0, variance roughly 1) at each step during training, keeping values from drifting too large or too small as they pass through many layers. Layer normalization specifically is used throughout transformers (see Attention & Transformers).
- Residual (skip) connections: instead of a layer's output being
purely a function of its input, add the input back to the output
(
output = layer(x) + x). This gives gradients a direct path backward that skips the layer's local gradient entirely, so even if a particular layer's local gradient is small, the gradient signal isn't lost — it can flow through the skip connection instead. This one architectural trick is a major reason networks with dozens or hundreds of layers (again, transformers included) are trainable at all.
The backward recursion above is a repeated matrix product. Unrolling it across layers, the gradient reaching layer carries the factor:
Each is a Jacobian — the derivative of one layer's output with respect to its input. The magnitude of the whole product is governed roughly by the product of their singular values. If the typical scale factor per layer is , the gradient at depth scales like :
- over 50 layers gives — vanishing
- over 50 layers gives — exploding
Because the dependence is exponential in depth, there is no safe middle ground you can tune your way into; has to be held near 1 by construction. That is exactly what the three fixes above do. Initialization schemes set so the initial scale is near 1, normalization layers rescale activations back to a fixed variance at every step, and a residual connection changes the Jacobian to , whose product stays well-behaved because the identity term guarantees a path with scale exactly 1.
Recap
- A neuron: weighted sum + bias, then a nonlinearity.
- Forward pass: chain these operations to produce a prediction and a loss.
- Backward pass (backprop): apply the chain rule in reverse through the computation graph to get the exact gradient of the loss with respect to every parameter, in one efficient pass.
- Gradient descent (previous lesson) then uses those gradients to update every parameter.
Every architecture in the rest of this course — CNNsCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently., transformers — is still trained by exactly this mechanism. What changes is only the shape of the computation graph.