ViT (Vision Transformer)
ViT applies the transformer's self-attention mechanism directly to images, splitting them into patches treated like sequence tokens, without CNN-style locality assumptions.
The Vision Transformer (ViT), introduced in 2020, applies self-attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. directly to images — splitting an image into fixed-size patches and treating each like a token in a sequence — instead of using CNNCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently.-style convolutions. Given enough training data, ViT matches or beats CNNs despite having no locality assumption built in; the model learns spatial relationships purely from data. It's evidence that attention is a genuinely general-purpose mechanism, not a language-specific trick.
How it works
The image is cut into non-overlapping patches — 16x16 pixels in the
original model, giving 196 patches for a 224x224 input. Each patch is
flattened and pushed through a single linear projection to produce a
token embedding, a learned classification token is prepended, and a
positional encodingPositional EncodingPositional encoding injects word-order information into a transformer, since self-attention alone treats a sequence's tokens as an unordered set. is added, since
attention on its own has no notion of where a patch sat in the grid.
From that point it is a standard transformer encoder: alternating
self-attention and MLP blocks with residual connections and layer
normalization. The classification head reads the final representation
of the prepended token. That patch projection is the only
image-specific component in the entire architecture.
When it breaks
- Trained from scratch on ImageNet-scale data alone, ViT underperforms a comparable CNN. It only pulls ahead after pretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability. on a much larger corpus, since it has to learn from data the locality that convolution gets for free.
- Attention cost grows quadratically with patch count, so halving the patch size roughly quadruples compute — high-resolution inputs get expensive quickly.
- Positional encodings are tied to the training resolution; running at a different input size requires interpolating them.
- The training recipe is load-bearing rather than incidental. Strong augmentationData AugmentationData augmentation expands a training set by applying label-preserving transformations (crops, flips, color jitter) to existing examples, fighting overfitting for free., long schedules, and careful optimizer and warmup settings are all required to reproduce reported results.
See also: AttentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer., CNNCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently., TransformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.
Learn more: Computer Vision · Paper: An Image Is Worth 16x16 Words
Mentioned in
Lessons where this comes up in context.
ResNet
ResNet is a deep CNN architecture that introduced residual (skip) connections, enabling networks with 50-150+ layers to train successfully.
CLIP (Contrastive Language-Image Pretraining)
CLIP trains an image encoder and a text encoder together with a contrastive loss so that matching image-caption pairs land close together in one shared embedding space.