Residual Connection
A residual (skip) connection adds a layer's input back to its output, giving gradients a direct path backward and making very deep networks trainable.
A residual connection (or "skip connection") adds a layer's input
back to its output — output = layer(x) + x — instead of the output
being purely a function of the input. This gives gradients a direct path
backward during backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network. that
bypasses the layer's own local gradient, preventing the signal from being
lost even when that gradient is small. Introduced prominently in
ResNetResNetResNet is a deep CNN architecture that introduced residual (skip) connections, enabling networks with 50-150+ layers to train successfully., residual connections are what make networks
with dozens or hundreds of layers — including
transformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. — trainable at all.
How it works
Because the derivative of x with respect to x is 1, the gradient
arriving at a residual block splits: part flows through layer(x), and
part passes straight through unchanged. Across N stacked blocks the
gradient reaching the first layer is a sum of paths rather than a
single long product, so it cannot decay geometrically with depth.
A second framing: the block only has to learn the residual
layer(x) = target - x. If the identity is already a good answer, the
block can learn to output near-zero, so adding layers never has to make
the network worse.
In practice the block is paired with normalization. Transformers use the
pre-norm arrangement x + sublayer(norm(x)), which keeps the skip
path free of any normalization and trains more stably than the original
post-norm form.
When it breaks
- Shape mismatch on the skip. When a block changes channel count or
spatial resolution,
xcannot be added directly; a1x1convolution or linear projection on the shortcut is required, as in ResNetResNetResNet is a deep CNN architecture that introduced residual (skip) connections, enabling networks with 50-150+ layers to train successfully.'s downsampling blocks. - Activation growth. Repeated addition makes the residual stream's
magnitude grow with depth. Without normalization or scaled
initializationWeight InitializationWeight initialization sets a neural network's starting parameter values, using schemes like Xavier/He initialization to keep gradients well-scaled from the first training step., very deep stacks
drift into numerical trouble, especially in
float16. - Post-norm instability. Placing normalization after the addition puts it on the skip path and typically requires a learning-rate warmup to train at all.
- Not a substitute for capacity. Skips fix optimization, not underfitting; a too-narrow network stays too narrow.
See also: ResNetResNetResNet is a deep CNN architecture that introduced residual (skip) connections, enabling networks with 50-150+ layers to train successfully., Vanishing/Exploding GradientsVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training.
Learn more: Neural Networks & Backprop · Computer Vision
Mentioned in
Lessons where this comes up in context.
Batch/Layer Normalization
Batch and layer normalization rescale a layer's outputs during training to keep values well-behaved as they pass through many stacked layers.
Universal Approximation Theorem
The universal approximation theorem proves that a feedforward network with even one hidden layer can approximate any continuous function, given enough neurons.