Architectures

Computer Vision

Convolutions, pooling, CNNs, transfer learning

The neural networks from the last lesson treat their input as a flat list of numbers. That's a bad fit for images: a photo is a grid of pixels where nearby pixels are related, and a flat feedforward layer throws that spatial structure away — it would need a separate weight for every pixel position, with no way to recognize "this is an edge" regardless of where in the image the edge appears. Convolutional neural networks (CNNs) fix this by building spatial structure into the architecture itself.

The whole architecture is one pipeline repeated a few times, so it is worth having the shape of it in mind before the pieces:

Each stage below turns one box into the next: raw pixels become feature maps, feature maps get downsampled and stacked into higher-level features, and what's left gets flattened into a single classification decision.

The convolution operation

A convolution slides a small grid of learnable numbers — a kernel (or filter), often 3×3 or 5×5 — across the image, and at each position computes a weighted sum of the pixels it currently overlaps:

output[i, j] = sum over (di, dj) of kernel[di, dj] * image[i + di, j + dj]

The kernel's weights are learned, exactly like any other parameter, via gradient descent and backprop. What makes this powerful is what it implies:

  • Local connectivity: each output value depends only on a small neighborhood of the input, not the whole image — matching the intuition that whether a pixel is part of an edge depends on its immediate neighbors, not on a pixel in the opposite corner.
  • Parameter sharing: the same kernel is applied at every position in the image. A kernel that learns to detect vertical edges detects them wherever they appear — top-left or bottom-right — using one small set of weights instead of one per pixel. This is what lets CNNs generalize across position and train with vastly fewer parameters than a fully connected layer would need for the same input size.

The output of applying a kernel across an image is called a feature map — a new grid where each value indicates how strongly that pattern (whatever the kernel learned to detect) appeared at that location.

Slide a few classic hand-designed kernels over a test image below and watch the feature map build up, cell by cell, exactly as the sliding window described above:

Input image (10×10)

Kernel (3×3)

-1
0
1
-1
0
1
-1
0
1

Feature map (8×8)

0 / 64

A real convolutional layer learns many kernels in parallel (commonly dozens to hundreds), each producing its own feature map — one might learn to detect vertical edges, another horizontal edges, another a particular color transition — stacked together into a multi-channel output that the next layer takes as input.

Conv layer vs. dense layer, same image

Take a 224×224 RGB image and a layer producing 64 feature maps at the same resolution.

As a convolution, 3×3 kernels over 3 input channels: 3 × 3 × 3 × 64 ≈ 1,700 weights (plus 64 biases).

As a fully connected layer over the same input and output: (224 × 224 × 3) × (224 × 224 × 64) ≈ 150,000 × 3,200,000 ≈ 5 × 10¹¹ weights — roughly half a trillion, around 2 TB in 32-bit floats.

That is a factor of about 10⁸, for a layer computing something broadly similar. And the conv layer's count does not depend on image size at all: feed it a 1024×1024 photo and it still has 1,700 weights, while the dense version grows with the fourth power of resolution.

Strides and padding

Two more knobs control exactly how a convolution scans the image:

  • Stride: how many pixels the kernel moves between positions. Stride 1 (the default assumed above) produces a feature map close to the input's size; stride 2 skips every other position, halving the output's spatial size directly — an alternative to pooling for downsampling.
  • Padding: convolutions naturally shrink the output slightly, since a kernel can't center on a pixel too close to the edge without going out of bounds. Adding a border of zero-valued pixels around the image ("padding") lets the output stay the same size as the input, if desired.

The output size of a convolution follows directly from these: output_size = (input_size - kernel_size + 2 * padding) / stride + 1. You won't compute this by hand in practice — frameworks handle it — but it's worth being able to sanity-check why a feature map came out the size it did.

Pooling

After a convolution, networks typically apply pooling — most commonly max pooling: divide the feature map into small regions (e.g. 2×2) and keep only the maximum value from each. This does two things:

  1. Downsamples: shrinks the spatial size of the data as it moves deeper into the network, keeping compute manageable.
  2. Adds translation tolerance: if a detected feature shifts by a pixel or two, max pooling is likely to still pick it up, since it only cares about the strongest activation in a neighborhood, not its exact position.

Stacking into a CNN

A typical CNN alternates convolution and pooling layers, then finishes with one or more regular (fully connected) layers to produce a final prediction:

image → [conv → activation → pool] × N → flatten → fully-connected → output

The key architectural idea is hierarchical feature learning: early layers (closest to the raw pixels) learn simple, generic features — edges, color blobs, simple textures. Each successive layer combines the previous layer's features into more complex ones — corners and simple shapes in the middle layers, then object parts (eyes, wheels, fur texture) in later layers, and finally whole-object representations right before the output. Nobody hand-designs this hierarchy — it emerges from training, which is exactly the feature-engineering bottleneck from History & Landscape being solved automatically.

Named architectures: a brief tour

A handful of specific CNN architectures mark the field's progress and are worth recognizing by name:

  • LeNet (1998, Yann LeCun): the original convolution-plus-pooling design, applied to handwritten digit recognition — small by modern standards, but the architectural template every later CNN follows.
  • AlexNet (2012): the ImageNet-winning network from History & Landscape — essentially a scaled-up LeNet trained on GPUs with ReLU activations, proving deep CNNs could dramatically outperform hand-engineered feature pipelines.
  • VGG (2014): showed that simply stacking many small (3×3) convolution layers very deep — 16 to 19 layers — outperformed shallower networks with larger kernels, reinforcing that depth itself was valuable.
  • ResNet (2015): pushed depth much further (50, 101, even 152+ layers) by introducing residual connections — the skip-connection fix from Neural Networks & Backprop that keeps gradients flowing through very deep networks. Before ResNet, networks much deeper than VGG actually got worse, not better — a direct symptom of the vanishing-gradient problem; residual connections were the fix that made extreme depth practical, in vision first and transformers later.

Each of these solved a specific obstacle the previous one hit, which is the same pattern seen throughout this course: architecture progress is mostly a series of fixes to specific training failure modes, not unrelated leaps.

Beyond classification

Everything above describes image classification — one label per image. Two other common vision tasks build on the same convolutional feature extractor but change what the final layers do with it:

  • Object detection: output not just a label but a bounding box (location) for each object in an image, potentially many per image.
  • Semantic/instance segmentation: classify every pixel in the image, producing a full mask of which pixels belong to which object rather than a single label or box.

Both reuse the same hierarchical features a classifier learns — the difference is entirely in the output layers and loss function, another instance of the "model, loss, optimize" recipe from ML Fundamentals being reused with a different head bolted onto the same backbone.

Data augmentation

A cheap way to fight overfitting (see ML Fundamentals) specifically for images: artificially expand the training set by applying label-preserving transformations — random crops, flips, rotations, color jitter — to each training image. A photo of a cat flipped horizontally is still obviously a cat, so this manufactures new, valid training examples for free, and it's standard practice in nearly every CNN training pipeline.

Transfer learning

Training a large CNN from scratch needs a huge labeled dataset (ImageNet has over a million labeled images) and significant compute. Transfer learning sidesteps this: take a network already trained on a large general dataset, and reuse it for a new, related task.

The intuition: the early and middle layers of a trained CNN have already learned generic visual features (edges, textures, shapes) that are useful for almost any vision task, not just the one the network was originally trained on. Two common strategies:

  • Feature extraction: freeze the pretrained network's weights entirely, and train only a new final layer (or few layers) on your smaller, task-specific dataset. Fast, and works well when your data is limited.
  • Fine-tuning: start from the pretrained weights, but continue training some or all of the network (usually with a smaller learning rate) on your new data — lets the model adapt its features more specifically to the new task, at the cost of needing more data and compute than pure feature extraction.

Transfer learning is why a team with a few thousand labeled photos of, say, manufacturing defects can still get a strong classifier: they aren't learning "what an edge looks like" from scratch, only reusing that from a network pretrained on a much larger, unrelated dataset and adapting the last stage to their specific categories.

Recap and what's next

CNNs solve vision by hard-coding two structural assumptions into the architecture — locality and translation invariance — that turn out to match how images are actually structured. That's a general pattern worth remembering: good architectures encode assumptions about the data's structure — an inductive bias — so the network doesn't have to learn those assumptions from scratch.

Why do Vision Transformers need more training data than CNNs?

Worth flagging before moving on: this hard-coded assumption isn't the final word. Vision Transformers (ViT), introduced in 2020, apply the attention mechanism from the next lesson directly to images — splitting an image into patches and treating them like tokens in a sequence — and match or beat CNNs given enough training data, without locality built in at all; the model learns spatial relationships from data instead. This is a preview of a broader theme: attention has proven to be a genuinely general-purpose mechanism, not just a language-specific trick.

CNNVision Transformer
Inductive biasLocality and translation equivariance built inAlmost none; learned from data
Data hungerWorks from modest datasetsNeeds large-scale pretraining to match CNNs
Receptive fieldLocal at first, global only with depthGlobal from the first layer

The next lesson looks at a different structural problem — sequences, where order and long-range relationships matter — and the architecture (attention) built to handle it.

On this page