Deep Residual Learning for Image Recognition
He et al., 2015 — solved the mystery of why stacking more layers made deep networks worse, not better, by adding a shortcut connection around every block. Won ImageNet 2015 and made 100+ layer networks routine.
"Deep Residual Learning for Image Recognition" (Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, 2015) introduced ResNetResNetResNet is a deep CNN architecture that introduced residual (skip) connections, enabling networks with 50-150+ layers to train successfully. — the architecture change that made it possible to train neural networks dozens or hundreds of layers deep, something that had been quietly breaking every deeper network before it.
What problem it solved
Intuition says a deeper network should never do worse than a shallower one — in the worst case, the extra layers could just learn to pass their input through unchanged. In practice, researchers found the opposite: past a certain depth, adding more layers made both training and test accuracy worse, and not because of overfitting — the deeper network's own training error was higher. Something about the way layers were stacked made it genuinely hard for a deep network to learn even the trivial "do nothing" function, let alone something useful.
The key idea
Instead of asking each block of layers to learn a full transformation of its input, restructure the block to learn a residual — the difference between its input and its desired output — and add that residual back onto the input via a skip connection:
output = F(x) + xIf the ideal function really is close to "do nothing," the block only has to push toward zero, which is a far easier optimization target than learning an identity mapping from scratch through several nonlinear layers. The skip connection also gives gradients a direct, unobstructed path backward through the network during backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network., which is why residual networks train far more reliably at extreme depth than plain stacks of layers do.
Why it mattered
ResNet won the ImageNet 2015 classification competition and, more importantly, made "just add more layers" a viable strategy again — the paper's own experiments went as deep as 152 layers, roughly 8× deeper than the VGG networks that came before it, with lower error. The residual connection itself outlived the specific architecture: it's now a default building block used far outside computer vision, including inside the transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. blocks that every modern LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. is built from.
Authors: Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun (Microsoft Research)
Read the paper: arXiv:1512.03385
Learn more: ResNetResNetResNet is a deep CNN architecture that introduced residual (skip) connections, enabling networks with 50-150+ layers to train successfully. · Computer Vision
ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, Sutskever & Hinton, 2012 — the paper widely credited with restarting deep learning, by winning ImageNet with a deep CNN trained on GPUs by a margin nobody expected.
An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy et al., 2020 — showed that a plain transformer, with no convolutions at all, could match or beat CNNs on image classification if trained on enough data, by treating an image as a sequence of patches.