Encoder-Decoder Architecture
Encoder-only, decoder-only, and encoder-decoder describe which half of the original transformer a model uses, determining whether it's built for understanding, generation, or both.
The original transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. paper describes two halves, and different model families use different combinations: encoder-only (e.g. BERT) uses unmasked attention and is good at understanding fixed text; decoder-only (e.g. GPT, most current LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama.) uses causal maskingCausal MaskingCausal masking restricts self-attention so each token can only attend to earlier positions, not future ones — required for training and running autoregressive generation. exclusively and is built for generation; encoder-decoder (e.g. T5) pairs an unmasked encoder over the input with a causal decoder that also attends back to the encoder's output — common for translation and summarization.
How it works
The difference is entirely in which positions may attend to which.
- Encoder-only: bidirectional attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer., every token sees the whole input. Trained with masked-token objectives, it yields one contextual vector per position and is used for classification, tagging, and embeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space. extraction.
- Decoder-only: each token sees only its predecessors, so one forward pass supervises every position at once and the same graph serves both training and generation.
- Encoder-decoder: the encoder reads the source bidirectionally; the decoder is causal over the target and adds a cross-attention sublayer whose Queries come from the decoder while Keys and Values come from the encoder output. The source is read in full while the target is produced one token at a time.
When it breaks
- Encoder-only models can't generate. Prompting BERT for free text is a category error — there is no next-token head to sample from.
- Cross-attention doubles the plumbing. Encoder-decoder inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. keeps two caches, the encoder output computed once and the decoder's own KV cacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable.; confusing them is a common source of wrong outputs on batched inputs.
- Decoder-only embeddings are lopsided. Under causal attention the last token has seen everything and the first has seen nothing, so naive mean-pooling for retrieval underperforms an encoder.
See also: TransformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs., Causal MaskingCausal MaskingCausal masking restricts self-attention so each token can only attend to earlier positions, not future ones — required for training and running autoregressive generation.
Learn more: Attention & Transformers
Causal Masking
Causal masking restricts self-attention so each token can only attend to earlier positions, not future ones — required for training and running autoregressive generation.
RNN (Recurrent Neural Network)
An RNN is a neural network for sequences that processes input one step at a time, carrying a hidden state forward — the main predecessor to transformers.