ResNet
ResNet is a deep CNN architecture that introduced residual (skip) connections, enabling networks with 50-150+ layers to train successfully.
ResNet (2015) is a CNNCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently. architecture that pushed
network depth far beyond what was previously trainable (50, 101, even
150+ layers) by introducing residual connections: adding a layer's
input back to its output (output = layer(x) + x), giving gradients a
direct path backward that skips the layer's local gradient. Before
ResNet, networks much deeper than prior architectures like VGG actually
performed worse — a symptom of vanishing gradients. Residual
connections are now used pervasively, including in
transformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs..
How it works
A residual block computes y = F(x) + x, where F is two or three
convolutions with normalization and a non-linearity between them. The
addition is the whole trick: the identity path has gradient exactly 1,
so whatever gradient reaches a block's output also reaches its input
undiminished, however many blocks sit above it. Signal stops shrinking
multiplicatively with depth, which is what defeats
vanishing gradientsVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training.. It also changes
what the layers are asked to learn: F only has to model the
difference from the identity, so an unhelpful block can settle near
zero instead of learning identity from scratch. Deeper variants use a
bottleneck block — 1x1 reduce, 3x3, 1x1 expand — with
batch normalizationBatch/Layer NormalizationBatch and layer normalization rescale a layer's outputs during training to keep values well-behaved as they pass through many stacked layers. after each
convolution.
When it breaks
- When a block changes channel count or spatial size the shapes stop
matching and the shortcut needs a projection (
1x1convolution with stride). Getting this wrong is the single most common implementation bug. - Depth still has diminishing returns — 1000-layer variants gain little over ResNet-152 and overfit smaller datasets.
- Ordering matters: placing normalization and activation before the convolutions (pre-activation) trains very deep stacks more reliably than the original post-activation arrangement.
- Skip connections raise activation memory during training, since each block's input must stay live until the addition.
See also: CNNCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently., AlexNetAlexNetAlexNet is the 2012 deep convolutional network that won the ImageNet competition by a wide margin, sparking the deep learning boom., BackpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network.
Learn more: Computer Vision · Paper: Deep Residual Learning for Image Recognition
Mentioned in
Lessons where this comes up in context.
AlexNet
AlexNet is the 2012 deep convolutional network that won the ImageNet competition by a wide margin, sparking the deep learning boom.
ViT (Vision Transformer)
ViT applies the transformer's self-attention mechanism directly to images, splitting them into patches treated like sequence tokens, without CNN-style locality assumptions.