Foundations

Neural Networks & Backprop

From Karpathy's micrograd approach — building a tiny neural net and stepping through forward/backward passes

ML Fundamentals established the loop: compute a loss, compute its gradient with respect to the parameters, step against the gradient. This lesson is about the one piece left unexplained — how you actually compute that gradient when your model is a neural network with millions of parameters chained through dozens of layers. The answer is an algorithm called backpropagation, and understanding it from scratch (the micrograd approach) demystifies everything that follows.

What a neuron actually is

A single artificial neuron is almost embarrassingly simple: it takes some inputs, multiplies each by a weight, adds them up with a bias term, and passes the result through a nonlinear function:

z = w1*x1 + w2*x2 + ... + wn*xn + b
a = activation(z)

The weights (w1...wn) and bias (b) are the learnable parameters. The activation function (commonly ReLU(z) = max(0, z), or historically tanh / sigmoid) introduces nonlinearity — without it, stacking layers of neurons would collapse into one big linear function, no more expressive than a single layer. Nonlinearity is what lets networks represent curved decision boundaries and complex functions at all.

A layer is a bunch of these neurons run in parallel on the same input. A network is layers stacked so each layer's output feeds the next layer's input. That's the entire architecture of a basic feedforward neural network, also called a multi-layer perceptron (MLP) — the complexity in modern AI comes from what's connected to what and how (see Computer Vision and Attention & Transformers), not from the neuron itself.

Why any of this can represent complex functions

A natural question: why does chaining together these simple weighted-sum-plus-nonlinearity units let you represent essentially anything? The universal approximation theorem answers this formally: a feedforward network with even a single hidden layer, given enough neurons in that layer, can approximate any continuous function to arbitrary precision. This is a real, proven result — not marketing language.

The catch is that "enough neurons" can be an impractically large number for a single layer, and the theorem says nothing about whether such a network is learnable via gradient descent in practice. This is exactly why depth matters more than the theorem alone would suggest: a deeper network can represent the same function with far fewer total parameters, and turns out to be far more trainable in practice — see "Where the 'deep' in deep learning comes from," below.

Choosing an activation function

Not all nonlinearities are equal in practice, and the choice has a real history behind it:

  • Sigmoid (1 / (1 + e^-z)) and tanh were the original defaults — smooth, bounded outputs. Their flaw: for large positive or negative z, the function flattens out almost completely, so its derivative is close to zero. Chained across many layers (see vanishing gradients, below), this made deep networks with these activations extremely slow or impossible to train.
  • ReLU (max(0, z)) fixed this for positive inputs — its derivative is exactly 1 wherever z > 0, so it doesn't shrink gradients flowing through active neurons. It's cheap to compute and was a major factor in making deep networks trainable at all. Its own flaw: a neuron whose input is always negative outputs zero forever and stops learning (the "dying ReLU" problem).
  • GELU and other smooth ReLU variants (used in most modern transformers, including the ones in Attention & Transformers) soften ReLU's sharp corner at zero, which empirically trains slightly better at the scale modern models operate at, at a small extra compute cost.
ActivationDerivative behaviorMain failure mode
SigmoidNear zero for large inputsVanishing gradients with depth
TanhNear zero for large inputsVanishing gradients with depth
ReLUExactly 1 when z > 0, else 0Dying neurons stuck at zero
GELUSmooth, nonzero near zeroSlightly more compute per unit

The forward pass

Running an input through the network to produce an output is called the forward pass. Concretely, for a tiny 2-layer network computing a prediction y_pred from an input x:

h  = ReLU(W1 · x + b1)      # hidden layer
y_pred = W2 · h + b2        # output layer
loss = (y_pred - y_true)^2  # e.g. squared error against the true label

Every one of those operations — multiply, add, ReLU, subtract, square — is a simple, differentiable function. That fact is the entire key to what comes next.

The same graph is traversed twice: values flow left to right to produce a loss, then gradients flow right to left back to the parameters.

The chain rule, and why it scales

To do gradient descent we need ∂loss/∂W1, ∂loss/∂b1, ∂loss/∂W2, ∂loss/∂b2 — how the loss changes with respect to every parameter, even ones buried several operations deep. Computing each of these from scratch, independently, would be enormously wasteful — many of them share intermediate work.

Backpropagation is just the chain rule from calculus, applied systematically: if loss depends on y_pred which depends on h which depends on W1, then

∂loss/∂W1 = (∂loss/∂y_pred) · (∂y_pred/∂h) · (∂h/∂W1)

The trick is order of computation: work backward from the loss, computing each local derivative once and reusing it for everything upstream. That's why it's called backpropagation — you build a computation graph forward, then walk it backward, accumulating gradients.

Stepping through it, node by node

This is the exercise worth actually doing once (as Karpathy's micrograd does): build a tiny scalar "autograd" engine where every number remembers (a) the operation that produced it and (b) its inputs. Two operations, concretely:

# forward: c = a * b
# backward, given dL/dc (gradient flowing in from later in the graph):
#   dL/da = dL/dc * b
#   dL/db = dL/dc * a

# forward: c = a + b
# backward, given dL/dc:
#   dL/da = dL/dc * 1
#   dL/db = dL/dc * 1

Every operation you'll ever use (multiply, add, ReLU, exp, division, matrix multiply for real networks) has a similarly simple local gradient rule. Backprop is just: run the forward pass while recording the graph, then walk it in reverse, multiplying local gradients along each path and summing where paths merge (a value used twice contributes to the gradient twice). At the end, every parameter in the network has an exact gradient of the loss with respect to it — computed in roughly the same amount of work as the forward pass itself, regardless of how many parameters there are.

That last property — the gradient of all parameters costs about the same as one forward pass — is what makes training networks with billions of parameters computationally feasible at all. Without it, gradient descent at LLM scale would be a non-starter.

Backprop vs. nudging each parameter by hand

You could estimate gradients without any calculus: nudge one parameter by a tiny amount, re-run the network, and see how much the loss changed. That's a finite difference, and it needs one forward pass per parameter.

For a small MLP with 1M parameters:

  • Finite differences: 1,000,000 forward passes per gradient step
  • Backprop: 1 forward + 1 backward ≈ 3 forward passes' worth of work

That's a factor of roughly 300,000 — and the gap grows linearly with parameter count, so at a billion parameters it is a factor of hundreds of millions. Backprop is not a speedup over the naive method; it is the difference between possible and impossible.

The price is memory. Take a 10-layer MLP with 1,024 units per layer, batch size 64, stored as 4-byte floats: each layer's activations are 64 × 1024 × 4 ≈ 0.26 MB, so all ten together are only ~2.6 MB. Tiny here — but the same accounting at transformer scale is why training a model needs far more memory than running it.

A fully worked numeric example

Concrete numbers make this stick better than symbols alone. Take the smallest possible case: loss = (a * b + c), with a = 2, b = -3, c = 10.

Forward pass:

d = a * b   = 2 * -3  = -6
loss = d + c = -6 + 10 =  4

Backward pass, starting from dL/dL = 1 (the loss's gradient with respect to itself) and working backward using the local rules above:

dL/dd = dL/dloss * 1 = 1        # loss = d + c, local grad w.r.t. d is 1
dL/dc = dL/dloss * 1 = 1        # local grad w.r.t. c is 1

dL/da = dL/dd * b = 1 * -3 = -3 # d = a*b, local grad w.r.t. a is b
dL/db = dL/dd * a = 1 * 2  =  2 # local grad w.r.t. b is a

Read dL/da = -3 as: "increasing a by a small amount increases the loss by about 3 times that amount" — so gradient descent will decrease a.

Step through this exact example node by node, forward then backward — try changing a, b, or c and re-running it:

a2b-3×d = a×b?c10+loss = d+c?

Step 0 / 5: Start

Set a, b, c, then step through the forward pass, then backward.

Every gradient in a million-parameter network is computed by chaining exactly these two rules (multiply, add) — plus a handful of others for whatever operations appear in the graph — over and over, automatically. This is precisely what loss.backward() does in a real framework (see Tooling & The Dev Stack) — the manual version above is what autograd is doing under the hood.

Where the "deep" in deep learning comes from

"Depth" just means stacking many layers. Deeper networks can represent more complex functions with fewer total parameters than a single very wide layer — each layer builds on features the previous layer extracted, rather than trying to learn the whole mapping at once. Practically, depth introduces its own problems, and three specific fixes are worth knowing by name, since all three show up throughout this course:

  • Vanishing/exploding gradients: backprop multiplies local gradients along a chain (as in the worked example above). If those local gradients are consistently a bit less than 1, their product shrinks toward zero across many layers — early layers get almost no gradient signal and stop learning. If they're consistently a bit more than 1, the product explodes instead, and training diverges. This is the practical reason sigmoid/tanh fell out of favor for deep networks (see activation functions, above) — their small derivatives compound badly across depth.
  • Weight initialization: the starting values of the weights matter more than intuition suggests — initializing too large or too small makes vanishing/exploding gradients worse from step one. Schemes like Xavier/Glorot initialization and He initialization set the initial random weight scale based on the number of inputs/outputs of each layer, specifically to keep activations and gradients at a consistent scale as they pass through the network.
  • Normalization layers (batch normalization, layer normalization): rescale a layer's outputs (mean roughly 0, variance roughly 1) at each step during training, keeping values from drifting too large or too small as they pass through many layers. Layer normalization specifically is used throughout transformers (see Attention & Transformers).
  • Residual (skip) connections: instead of a layer's output being purely a function of its input, add the input back to the output (output = layer(x) + x). This gives gradients a direct path backward that skips the layer's local gradient entirely, so even if a particular layer's local gradient is small, the gradient signal isn't lost — it can flow through the skip connection instead. This one architectural trick is a major reason networks with dozens or hundreds of layers (again, transformers included) are trainable at all.

Recap

  • A neuron: weighted sum + bias, then a nonlinearity.
  • Forward pass: chain these operations to produce a prediction and a loss.
  • Backward pass (backprop): apply the chain rule in reverse through the computation graph to get the exact gradient of the loss with respect to every parameter, in one efficient pass.
  • Gradient descent (previous lesson) then uses those gradients to update every parameter.

Every architecture in the rest of this course — CNNs, transformers — is still trained by exactly this mechanism. What changes is only the shape of the computation graph.

On this page