Multimodal Models
How a single model handles more than one kind of input — CLIP's shared embedding space, how vision gets fed into a language model as tokens, and what actually breaks
Everything so far in this course has one modality at a time: Computer Vision processes images, LLMs process text. A multimodal model handles more than one — most commonly text and images — inside a single model, not by running two separate models and gluing their outputs together afterward. The distinction matters: gluing separate models together loses exactly the thing that makes multimodal models useful, which is a shared representation where a concept means the same thing regardless of which modality it arrived through.
CLIP: putting images and text in the same space
CLIPCLIP (Contrastive Language-Image Pretraining)CLIP trains an image encoder and a text encoder together with a contrastive loss so that matching image-caption pairs land close together in one shared embedding space. (Contrastive Language-Image Pretraining) is
the model that established how to get there. It trains two separate encoders — a
ViTViT (Vision Transformer)ViT applies the transformer's self-attention mechanism directly to images, splitting them into patches treated like sequence tokens, without CNN-style locality assumptions. for images, a transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. for text — but trains
them together, on one thing: a large set of (image, caption) pairs
scraped from the web.
The training signal is contrastive, the same underlying idea behind how word2vecword2vecword2vec (2013) was a technique for learning dense vector representations of words from raw text, a direct precursor to modern token embeddings. learns from context, generalized across modalities:
Take a batch of (image, caption) pairs. Encode every image and every
caption into a vector with its respective encoder.
For each real pair, the image vector and its true caption's vector should end up close together (high cosine similarity).
For every mismatched pairing within the batch — this image with some other caption in the batch — the two vectors should end up far apart.
Train both encoders jointly to make that true across millions of pairs, and the two encoders converge on a shared embeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space. space: a region where "a photo of a dog" and an actual photo of a dog land near each other, regardless of which encoder produced the vector.
Once trained, that shared space is directly useful on its own — zero-shot image classification, for instance, is just embedding an image and a set of candidate text labels, then picking whichever label's embedding is closest — no classification-specific training required. But its bigger role is as the building block for the next step.
Feeding vision into a language model
A shared embedding space doesn't by itself let a model have a conversation about an image. The dominant current approach connects a frozen (or lightly tuned) vision encoder to an existing LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama., so the LLM's already-trained language ability transfers over almost for free:
- Split the image into patches with a ViT-style encoder, exactly as Computer Vision describes.
- Run a small trainable projection layer that maps each patch embedding from the vision encoder's space into the LLM's own embedding space — the same dimensionality its text token embeddings live in.
- Feed those projected patch embeddings into the LLM's context window as if they were ordinary tokens, interleaved with the actual text tokens of the prompt.
From the transformer's self-attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. mechanism's point of view, an image patch and a word are the same kind of thing once this projection has run: a vector in context that every other position can attend to. The LLM was never retrained to fundamentally understand images — it's attending over a sequence that happens to contain some vectors sourced from pixels instead of tokenizationTokenizationTokenization converts raw text into a sequence of integers a model can process, typically via subword schemes like byte-pair encoding (BPE)..
A ViT typically splits an image into patches — a common configuration
uses roughly 16×16-pixel patches, so a modest 448×448 image becomes
roughly (448/16)² = 784 patches, and each patch becomes one token in
the LLM's context window. A single image can cost several hundred to
over a thousand tokens of context before the model has read a single
word of the actual prompt — which is why multimodal APIs typically
charge for images by an equivalent token count, and why high-resolution
images get downsampled or tiled rather than fed in at full resolution by
default.
Generation runs the other direction
Everything above is understanding — turning an image into something a language model can reason about and describe. Generating a new image from text is a different architecture entirely: Generative Models covers diffusion modelsDiffusion ModelA diffusion model learns to reverse a fixed process of gradually adding noise to data, generating new samples by denoising pure noise step by step., which are conditioned on a text embedding but otherwise don't share machinery with CLIP or a vision-language LLM. It's worth keeping the two directions separate in your head — "model that can describe a photo" and "model that can produce one from a caption" are built differently, even though both are routinely called multimodal.
Audio follows the same pattern
Audio-capable models generally reuse the same recipe: an audio encoder turns a waveform into a sequence of embeddings (frames of audio instead of image patches), a projection layer maps them into the LLM's space, and they're fed into context the same way. Nothing about the pattern is image-specific — it's a general recipe for bringing any modality that can be encoded into a sequence of vectors into a language model's context window, which is why new modalities tend to get added to existing LLMs faster than earlier architectures could support.
Where this actually breaks
- The modality gap. Even after contrastive training pulls matching pairs together, image embeddings and text embeddings in practice tend to cluster in two distinguishable regions of the shared space rather than fully overlapping — matching pairs are close relative to mismatched pairs, which is all contrastive training actually optimizes for, not perfect modality-agnostic overlap. Retrieval that crosses modalities (find the image for this caption) is measurably easier than the reverse in practice.
- Resolution is a real tradeoff, not just a cost one. Downsampling an image to control token cost throws away exactly the detail needed for tasks like reading small text in a photo or counting many small objects — a model can describe a scene well while failing a task that needs the fine detail that got downsampled away.
- Grounding failures compound with scale. A caption-heavy web dataset teaches strong associations between common objects and common captions, which is also what makes it easy for the model to describe a plausible scene that isn't the one in front of it — the training data rewardsRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one. plausibility, not verified grounding in the specific pixels given.