VAE (Variational Autoencoder)
A VAE encodes inputs into a distribution over a compact latent space rather than a fixed point, so sampling from that space and decoding produces plausible new outputs.
A VAE (Variational Autoencoder) compresses an input into a compact latent representation via an encoder, and reconstructs the original input from that latent vector via a decoder — similar in spirit to the embeddingsEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space. used elsewhere in this course. The "variational" part is what makes it generative rather than just compression.
How it works
Instead of encoding an input to a single fixed point, the encoder produces a distribution (see Probability DistributionProbability DistributionA probability distribution assigns a likelihood to each possible value of a random variable — the object every loss function is secretly built to measure the fit of.) over the latent space, and the model is trained so that sampling randomly from that latent space and decoding produces plausible new outputs — not just reconstructions of inputs it has already seen. Training optimizes a single, well-defined loss (no adversarial dynamics), which makes VAEs considerably more stable to train than a GANGAN (Generative Adversarial Network)A GAN trains a generator and a discriminator against each other, the generator learning to fool the discriminator into mistaking its output for real data..
When it breaks
- Historically blurrier outputs. The training objective tends to average over plausible reconstructions rather than commit to sharp detail, which is one reason diffusion modelsDiffusion ModelA diffusion model learns to reverse a fixed process of gradually adding noise to data, generating new samples by denoising pure noise step by step. became the dominant approach for high-fidelity image generation.
- The latent space's structure is a modeling assumption. Typically Gaussian, which is tractable but may not match the true shape of the data's underlying variation.
- Still useful as a compression stage, not just standalone. Modern diffusion models commonly run their denoising process in a compressed VAE latent space rather than on raw pixels — combining both techniques rather than choosing one.
See also: GANGAN (Generative Adversarial Network)A GAN trains a generator and a discriminator against each other, the generator learning to fool the discriminator into mistaking its output for real data., Diffusion ModelDiffusion ModelA diffusion model learns to reverse a fixed process of gradually adding noise to data, generating new samples by denoising pure noise step by step., EmbeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space.
Learn more: Generative Models
Mentioned in
Lessons where this comes up in context.
GAN (Generative Adversarial Network)
A GAN trains a generator and a discriminator against each other, the generator learning to fool the discriminator into mistaking its output for real data.
Diffusion Model
A diffusion model learns to reverse a fixed process of gradually adding noise to data, generating new samples by denoising pure noise step by step.