CNN (Convolutional Neural Network)
A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently.
A convolutional neural network (CNN) is a neural networkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. built around the convolution operation: a small learnable kernel slides across an image, computing weighted sums of local pixel neighborhoods. This encodes two structural assumptions well-suited to images — local connectivity and translation invariance — letting CNNs learn hierarchical visual features (edges → shapes → object parts) directly from pixels. Landmark CNNs include AlexNetAlexNetAlexNet is the 2012 deep convolutional network that won the ImageNet competition by a wide margin, sparking the deep learning boom. and ResNetResNetResNet is a deep CNN architecture that introduced residual (skip) connections, enabling networks with 50-150+ layers to train successfully..
How it works
A convolution layer holds a bank of kernels, each a small tensor of
shape (C_in, k, k). Every kernel slides across the input feature map
and produces one output channel, so a layer with C_out kernels maps
(C_in, H, W) to (C_out, H', W'). Weights are shared across all
spatial positions: a filter that detects a vertical edge detects it
anywhere in the frame, which is what makes CNNs far more
parameter-efficient than a dense layer over the same pixels. Stacking
layers widens the receptive field — each unit sees a larger patch of
the original image. Stride and pooling shrink the spatial dimensions as
channel count grows, with
batch normalizationBatch/Layer NormalizationBatch and layer normalization rescale a layer's outputs during training to keep values well-behaved as they pass through many stacked layers. and a
non-linearity between convolutions. Training is ordinary
backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network..
When it breaks
- Translation invariance is only approximate. Pooling and stride buy some tolerance, but rotation, scale, and viewpoint changes are not handled — which is why augmentation is effectively mandatory.
- The inductive bias that helps on modest datasets becomes a ceiling on very large ones, where ViTViT (Vision Transformer)ViT applies the transformer's self-attention mechanism directly to images, splitting them into patches treated like sequence tokens, without CNN-style locality assumptions. catches up and passes CNNs.
- Receptive fields grow slowly with depth, so relationships spanning opposite corners of an image need many layers, dilated convolutions, or attention to capture.
- Convolution assumes a regular grid, so it doesn't transfer to graphs, sets, or irregular point clouds without reformulation.
See also: AlexNetAlexNetAlexNet is the 2012 deep convolutional network that won the ImageNet competition by a wide margin, sparking the deep learning boom., ResNetResNetResNet is a deep CNN architecture that introduced residual (skip) connections, enabling networks with 50-150+ layers to train successfully., OpenCVOpenCVOpenCV is an open-source computer vision library providing classical (non-deep-learning) image processing algorithms, plus tooling to work with deep learning CV models.
Learn more: Computer Vision · Wikipedia: Convolutional neural network
Mentioned in
Lessons where this comes up in context.
- Applied & Agentic SystemsHow prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
- Attention & TransformersThe core architecture behind modern AI
- Computer VisionConvolutions, pooling, CNNs, transfer learning
- Generative ModelsGANs, VAEs, and diffusion models — how AI generates new images, audio, and video, as opposed to classifying or understanding existing content
- ML FundamentalsSupervised/unsupervised learning, loss functions, gradient descent
- Neural Networks & BackpropFrom Karpathy's micrograd approach — building a tiny neural net and stepping through forward/backward passes
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
MAP (Maximum a Posteriori Estimation)
MAP estimation extends MLE with a prior belief about the parameters themselves, via Bayes' theorem — and is precisely where weight-decay regularization comes from.
OpenCV
OpenCV is an open-source computer vision library providing classical (non-deep-learning) image processing algorithms, plus tooling to work with deep learning CV models.