Computer Vision
Convolutions, pooling, CNNs, transfer learning
The neural networksNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. from the last lesson treat their input as a flat list of numbers. That's a bad fit for images: a photo is a grid of pixels where nearby pixels are related, and a flat feedforward layer throws that spatial structure away — it would need a separate weight for every pixel position, with no way to recognize "this is an edge" regardless of where in the image the edge appears. Convolutional neural networksCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently. (CNNs) fix this by building spatial structure into the architecture itself.
The whole architecture is one pipeline repeated a few times, so it is worth having the shape of it in mind before the pieces:
Each stage below turns one box into the next: raw pixels become feature maps, feature maps get downsampled and stacked into higher-level features, and what's left gets flattened into a single classification decision.
The convolution operation
A convolution slides a small grid of learnable numbers — a kernel (or filter), often 3×3 or 5×5 — across the image, and at each position computes a weighted sum of the pixels it currently overlaps:
output[i, j] = sum over (di, dj) of kernel[di, dj] * image[i + di, j + dj]The kernel's weights are learned, exactly like any other parameter, via gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. and backprop. What makes this powerful is what it implies:
- Local connectivity: each output value depends only on a small neighborhood of the input, not the whole image — matching the intuition that whether a pixel is part of an edge depends on its immediate neighbors, not on a pixel in the opposite corner.
- Parameter sharing: the same kernel is applied at every position in the image. A kernel that learns to detect vertical edges detects them wherever they appear — top-left or bottom-right — using one small set of weights instead of one per pixel. This is what lets CNNs generalize across position and train with vastly fewer parameters than a fully connected layer would need for the same input size.
The output of applying a kernel across an image is called a feature map — a new grid where each value indicates how strongly that pattern (whatever the kernel learned to detect) appeared at that location.
A convolutional layer maps an input with channels to an output with channels. For a kernel, output channel at position is:
Three things fall out of this expression. The indices and appear only in , never in — that is parameter sharing written down: the same weights are used at every spatial position. The sums over and are bounded by , which is local connectivity. And the sum over runs over all input channels, so a kernel is never really — it is , and a layer holds weights plus biases regardless of how large the image is.
One pedantic note: this is technically cross-correlation, not
convolution, which would flip the kernel to .
Since is learned, the flip is irrelevant — the layer would just learn
the mirrored kernel — and every deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones. framework calls the
unflipped version conv.
Slide a few classic hand-designed kernels over a test image below and watch the feature map build up, cell by cell, exactly as the sliding window described above:
Input image (10×10)
Kernel (3×3)
Feature map (8×8)
A real convolutional layer learns many kernels in parallel (commonly dozens to hundreds), each producing its own feature map — one might learn to detect vertical edges, another horizontal edges, another a particular color transition — stacked together into a multi-channel output that the next layer takes as input.
Take a 224×224 RGB image and a layer producing 64 feature maps at the same resolution.
As a convolution, 3×3 kernels over 3 input channels: 3 × 3 × 3 × 64 ≈ 1,700 weights (plus 64 biases).
As a fully connected layer over the same input and output: (224 × 224 × 3) × (224 × 224 × 64) ≈ 150,000 × 3,200,000 ≈ 5 × 10¹¹ weights — roughly half a trillion, around 2 TB in 32-bit floats.
That is a factor of about 10⁸, for a layer computing something broadly similar. And the conv layer's count does not depend on image size at all: feed it a 1024×1024 photo and it still has 1,700 weights, while the dense version grows with the fourth power of resolution.
Strides and padding
Two more knobs control exactly how a convolution scans the image:
- Stride: how many pixels the kernel moves between positions. Stride 1 (the default assumed above) produces a feature map close to the input's size; stride 2 skips every other position, halving the output's spatial size directly — an alternative to pooling for downsampling.
- Padding: convolutions naturally shrink the output slightly, since a kernel can't center on a pixel too close to the edge without going out of bounds. Adding a border of zero-valued pixels around the image ("padding") lets the output stay the same size as the input, if desired.
The output size of a convolution follows directly from these:
output_size = (input_size - kernel_size + 2 * padding) / stride + 1.
You won't compute this by hand in practice — frameworks handle it — but
it's worth being able to sanity-check why a feature map came out the
size it did.
For input size , kernel size , padding and stride , the output size along each spatial dimension is:
The floor matters: when the stride doesn't divide evenly, the final window simply doesn't fit and gets dropped.
Setting and gives exactly — the "same" padding mode. For odd this is an integer, which is most of why kernel sizes are odd in practice, and with it costs a one-pixel border. That is the combination VGG-style stacks use throughout: spatial size is held fixed by the convolutions and changed only at the downsampling steps, so the architecture's resolution schedule stays easy to reason about.
With and the formula collapses to — the non-overlapping case used by 2×2 stride-2 max pooling, and, as it happens, by a Vision TransformerViT (Vision Transformer)ViT applies the transformer's self-attention mechanism directly to images, splitting them into patches treated like sequence tokens, without CNN-style locality assumptions.'s patch embeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space..
Pooling
After a convolution, networks typically apply pooling — most commonly max pooling: divide the feature map into small regions (e.g. 2×2) and keep only the maximum value from each. This does two things:
- Downsamples: shrinks the spatial size of the data as it moves deeper into the network, keeping compute manageable.
- Adds translation tolerance: if a detected feature shifts by a pixel or two, max pooling is likely to still pick it up, since it only cares about the strongest activation in a neighborhood, not its exact position.
Stacking into a CNN
A typical CNN alternates convolution and pooling layers, then finishes with one or more regular (fully connected) layers to produce a final prediction:
image → [conv → activation → pool] × N → flatten → fully-connected → outputThe key architectural idea is hierarchical feature learning: early layers (closest to the raw pixels) learn simple, generic features — edges, color blobs, simple textures. Each successive layer combines the previous layer's features into more complex ones — corners and simple shapes in the middle layers, then object parts (eyes, wheels, fur texture) in later layers, and finally whole-object representations right before the output. Nobody hand-designs this hierarchy — it emerges from training, which is exactly the feature-engineering bottleneck from History & Landscape being solved automatically.
The receptive field of a unit is the region of the original image that can influence it. For a stack of layers with kernel sizes and strides , it grows as:
With stride-1 3×3 convolutions every term is , so : the receptive field grows only linearly with depth, and 50 layers would still see just a 101-pixel window. Downsampling is what fixes this. Each stride-2 step doubles the product term, so every later layer's becomes , , — the growth turns geometric, and a handful of blocks is enough to cover an entire image.
So pooling and strided convolution are not only a compute optimisation. They are how a CNN reaches a global view at all, which is also why a single self-attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. layer is such a different proposition: its receptive field is the whole input from layer one.
Named architectures: a brief tour
A handful of specific CNN architectures mark the field's progress and are worth recognizing by name:
- LeNet (1998, Yann LeCunYann LeCunInvented the convolutional neural network in the late 1980s and deployed it commercially reading handwritten checks years before deep learning was fashionable.): the original convolution-plus-pooling design, applied to handwritten digit recognition — small by modern standards, but the architectural template every later CNN follows.
- AlexNetAlexNetAlexNet is the 2012 deep convolutional network that won the ImageNet competition by a wide margin, sparking the deep learning boom. (2012): the ImageNet-winning network from History & Landscape — essentially a scaled-up LeNet trained on GPUsGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires. with ReLU activations, proving deep CNNs could dramatically outperform hand-engineered feature pipelines.
- VGG (2014): showed that simply stacking many small (3×3) convolution layers very deep — 16 to 19 layers — outperformed shallower networks with larger kernels, reinforcing that depth itself was valuable.
- ResNetResNetResNet is a deep CNN architecture that introduced residual (skip) connections, enabling networks with 50-150+ layers to train successfully. (2015): pushed depth much further (50, 101, even 152+ layers) by introducing residual connectionsResidual ConnectionA residual (skip) connection adds a layer's input back to its output, giving gradients a direct path backward and making very deep networks trainable. — the skip-connection fix from Neural Networks & Backprop that keeps gradients flowing through very deep networks. Before ResNet, networks much deeper than VGG actually got worse, not better — a direct symptom of the vanishing-gradient problem; residual connections were the fix that made extreme depth practical, in vision first and transformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. later.
Each of these solved a specific obstacle the previous one hit, which is the same pattern seen throughout this course: architecture progress is mostly a series of fixes to specific training failure modes, not unrelated leaps.
Beyond classification
Everything above describes image classification — one label per image. Two other common vision tasks build on the same convolutional feature extractor but change what the final layers do with it:
- Object detection: output not just a label but a bounding box (location) for each object in an image, potentially many per image.
- Semantic/instance segmentation: classify every pixel in the image, producing a full mask of which pixels belong to which object rather than a single label or box.
Both reuse the same hierarchical features a classifier learns — the difference is entirely in the output layers and loss functionLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training., another instance of the "model, loss, optimize" recipe from ML Fundamentals being reused with a different head bolted onto the same backbone.
Data augmentation
A cheap way to fight overfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data. (see ML Fundamentals) specifically for images: artificially expand the training set by applying label-preserving transformations — random crops, flips, rotations, color jitter — to each training image. A photo of a cat flipped horizontally is still obviously a cat, so this manufactures new, valid training examples for free, and it's standard practice in nearly every CNN training pipeline.
Transfer learning
Training a large CNN from scratch needs a huge labeled dataset (ImageNet has over a million labeled images) and significant compute. Transfer learning sidesteps this: take a network already trained on a large general dataset, and reuse it for a new, related task.
The intuition: the early and middle layers of a trained CNN have already learned generic visual features (edges, textures, shapes) that are useful for almost any vision task, not just the one the network was originally trained on. Two common strategies:
- Feature extraction: freeze the pretrained network's weights entirely, and train only a new final layer (or few layers) on your smaller, task-specific dataset. Fast, and works well when your data is limited.
- Fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions.: start from the pretrained weights, but continue training some or all of the network (usually with a smaller learning rate) on your new data — lets the model adapt its features more specifically to the new task, at the cost of needing more data and compute than pure feature extraction.
Transfer learning is why a team with a few thousand labeled photos of, say, manufacturing defects can still get a strong classifier: they aren't learning "what an edge looks like" from scratch, only reusing that from a network pretrained on a much larger, unrelated dataset and adapting the last stage to their specific categories.
Recap and what's next
CNNs solve vision by hard-coding two structural assumptions into the architecture — locality and translation invariance — that turn out to match how images are actually structured. That's a general pattern worth remembering: good architectures encode assumptions about the data's structure — an inductive bias — so the network doesn't have to learn those assumptions from scratch.
Why do Vision Transformers need more training data than CNNs?
Worth flagging before moving on: this hard-coded assumption isn't the final word. Vision Transformers (ViT), introduced in 2020, apply the attention mechanism from the next lesson directly to images — splitting an image into patches and treating them like tokens in a sequence — and match or beat CNNs given enough training data, without locality built in at all; the model learns spatial relationships from data instead. This is a preview of a broader theme: attention has proven to be a genuinely general-purpose mechanism, not just a language-specific trick.
| CNN | Vision Transformer | |
|---|---|---|
| Inductive bias | Locality and translation equivariance built in | Almost none; learned from data |
| Data hunger | Works from modest datasets | Needs large-scale pretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability. to match CNNs |
| Receptive field | Local at first, global only with depth | Global from the first layer |
The next lesson looks at a different structural problem — sequences, where order and long-range relationships matter — and the architecture (attention) built to handle it.