AlexNet
AlexNet is the 2012 deep convolutional network that won the ImageNet competition by a wide margin, sparking the deep learning boom.
AlexNet is the deep CNNCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently. — built by Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton — that won the 2012 ImageNet competition by a wide margin, roughly halving the previous best error rate. Trained on GPUs with ReLU activations, it's widely credited as the moment that proved deep learning could dramatically outperform hand-engineered feature pipelines, kicking off the deep learning boom of the 2010s.
How it works
AlexNet stacks five convolutional layers followed by three fully-connected layers, ending in a 1000-way softmax over ImageNet classes — roughly 60 million parameters. Three choices mattered more than the raw depth:
- ReLU activationsActivation FunctionAn activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function. in place of
tanhor sigmoid, which trained substantially faster and didn't saturate. - GPU training. The network was split across two GTX 580 cards because 3 GB of memory couldn't hold it — an early, practical case of splitting a model across GPUsGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires..
- Dropout in the fully-connected layers plus heavy
augmentationData AugmentationData augmentation expands a training set by applying label-preserving transformations (crops, flips, color jitter) to existing examples, fighting overfitting for free.: random
224x224crops from256x256images, horizontal flips, and PCA-based color jitter, all aimed at overfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data..
When it breaks
- The large early kernels (
11x11, stride 4) and 4096-unit dense layers are enormously parameter-heavy; later architectures replaced them with stacks of3x3convolutions and global pooling. - Local response normalization, one of the paper's contributions, was later shown to contribute little and is effectively dead — batch normalization superseded it.
- Depth stops paying off at this scale. Naively stacking more layers in the AlexNet style makes accuracy worse, the problem residual connections were introduced to solve.
See also: CNNCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently., ResNetResNetResNet is a deep CNN architecture that introduced residual (skip) connections, enabling networks with 50-150+ layers to train successfully.
Learn more: History & Landscape · Computer Vision · Paper: ImageNet Classification with Deep Convolutional Neural Networks
Mentioned in
Lessons where this comes up in context.
OpenCV
OpenCV is an open-source computer vision library providing classical (non-deep-learning) image processing algorithms, plus tooling to work with deep learning CV models.
ResNet
ResNet is a deep CNN architecture that introduced residual (skip) connections, enabling networks with 50-150+ layers to train successfully.