Diffusion Model
A diffusion model learns to reverse a fixed process of gradually adding noise to data, generating new samples by denoising pure noise step by step.
A diffusion model generates new content by learning to reverse a process of gradually adding random noise to training data. It's the approach behind most current state-of-the-art image generators, and increasingly behind audio and video generation too.
How it works
The forward process is fixed arithmetic that needs no learning: take a real training example and progressively add small amounts of noise over many steps until it's indistinguishable from pure noise. Only the reverse direction is a network, trained to answer one narrow question at every noise level — what noise is in this input? — which reframes generation as a long sequence of simple, individually easy "remove a little noise" prediction problems, trained with an ordinary supervised loss functionLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. rather than an adversarial one.
Generation runs the trained network in reverse: start from random noise and repeatedly apply the denoising step, gradually turning noise into a coherent sample over many iterations — often dozens to hundreds of passes through the network.
When it breaks
- Slow to sample from. Unlike a GANGAN (Generative Adversarial Network)A GAN trains a generator and a discriminator against each other, the generator learning to fool the discriminator into mistaking its output for real data.'s single forward pass, generating one output requires running the network many times in sequence — the main cost diffusion models pay for their training stability.
- Usually run in a compressed latent space, not on raw data directly. Modern systems (e.g. Stable Diffusion) denoise inside a VAEVAE (Variational Autoencoder)A VAE encodes inputs into a distribution over a compact latent space rather than a fixed point, so sampling from that space and decoding produces plausible new outputs. latent space rather than raw pixels, since attention cost grows with the square of the number of positions being processed.
- Each step touches the whole output, not a region of it. It's easy to picture diffusion "painting in" an image gradually, but every step predicts noise across the entire input at once — coarse layout emerges early and fine texture late only because of which frequencies survive at each noise level, not because of spatial order.
See also: GANGAN (Generative Adversarial Network)A GAN trains a generator and a discriminator against each other, the generator learning to fool the discriminator into mistaking its output for real data., VAEVAE (Variational Autoencoder)A VAE encodes inputs into a distribution over a compact latent space rather than a fixed point, so sampling from that space and decoding produces plausible new outputs., TransformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.
Learn more: Generative Models
Mentioned in
Lessons where this comes up in context.
- Applied & Agentic SystemsHow prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
- Attention & TransformersThe core architecture behind modern AI
- Generative ModelsGANs, VAEs, and diffusion models — how AI generates new images, audio, and video, as opposed to classifying or understanding existing content
- Multimodal ModelsHow a single model handles more than one kind of input — CLIP's shared embedding space, how vision gets fed into a language model as tokens, and what actually breaks
VAE (Variational Autoencoder)
A VAE encodes inputs into a distribution over a compact latent space rather than a fixed point, so sampling from that space and decoding produces plausible new outputs.
Transformer
The transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.