ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, Sutskever & Hinton, 2012 — the paper widely credited with restarting deep learning, by winning ImageNet with a deep CNN trained on GPUs by a margin nobody expected.
"ImageNet Classification with Deep Convolutional Neural Networks" (Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton, 2012) introduced AlexNet — the single result most historians point to as the moment deep learning went from a niche, mostly-abandoned approach to the field's dominant paradigm.
What problem it solved
By 2012, neural networks had a credibility problem. The first AI winterMarvin MinskyCo-founder of the MIT AI Lab and a founding figure of the field — his 1969 book with Seymour Papert exposed the perceptron's limits and helped trigger the first AI winter. had been triggered partly by doubts about what networks like the perceptron could learn, and even after backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network. revived multi-layer networks in the 1980s, hand-engineered feature pipelines — features designed by domain experts rather than learned from data — still beat neural networks on most real computer vision benchmarks. The question this paper effectively answered: given enough data and enough compute, could a network that learns its own features actually outperform decades of hand-engineered ones?
The key idea
Scale up an existing idea — the deep CNNCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently. — further than anyone had successfully trained one before, and pair it with a handful of practical choices that made that scale trainable at all:
- ReLU activations instead of
tanh/sigmoid, which trained several times faster and didn't saturate the way earlier activation functions did. - GPU training. The full network didn't fit in one GPU's memory, so it was split across two, with careful engineering to keep them communicating efficiently — an early, hands-on case of the model-parallelism techniques large models still rely on.
- Dropout and heavy data augmentation (random crops, flips, color jitter) to fight overfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data. on a network with roughly 60 million parameters, large for the time.
None of these ideas were individually brand new — the contribution was combining them at a scale nobody had successfully made work before, on a dataset (ImageNet) finally large enough for that scale to pay off.
Why it mattered
AlexNet won the 2012 ImageNet competition by a margin that stunned the field — roughly halving the error rate of the next-best entry, which was still using hand-engineered features. Within a few years, virtually every serious computer vision system was built on deep CNNs, and the "more data, more compute, learned features beat hand-engineered ones" lesson generalized far past vision — it's the same bet that later drove scaling large language models. ResNet and ViT, covered elsewhere in this section, are both direct continuations of the architecture line AlexNet reopened.
Authors: Alex Krizhevsky, Ilya Sutskever, Geoffrey Hinton (University of Toronto)
Read the paper: NeurIPS 2012
See also: Alex KrizhevskyAlex KrizhevskyBuilt AlexNet in 2012 with Ilya Sutskever and Geoffrey Hinton, the deep convolutional network whose ImageNet win is widely credited with kicking off the deep learning boom. · Ilya SutskeverIlya SutskeverCo-authored AlexNet as a student, then sequence-to-sequence learning, then co-founded OpenAI and helped drive the bet that scale would produce GPT-level language models. · Geoffrey HintonGeoffrey HintonKnown as a godfather of deep learning — co-authored backpropagation in 1986, co-invented the Boltzmann machine and dropout, and co-authored AlexNet, the 2012 result that restarted the field.
Learn more: AlexNetAlexNetAlexNet is the 2012 deep convolutional network that won the ImageNet competition by a wide margin, sparking the deep learning boom. · Computer Vision · History & Landscape
Efficient Estimation of Word Representations in Vector Space
Mikolov et al., 2013 — showed that a shallow, cheaply-trained neural network could turn words into vectors that captured meaning, kicking off the embedding era that everything from search to LLMs now depends on.
Deep Residual Learning for Image Recognition
He et al., 2015 — solved the mystery of why stacking more layers made deep networks worse, not better, by adding a shortcut connection around every block. Won ImageNet 2015 and made 100+ layer networks routine.