RNN (Recurrent Neural Network)
An RNN is a neural network for sequences that processes input one step at a time, carrying a hidden state forward — the main predecessor to transformers.
A recurrent neural network (RNN) processes a sequence one element at a time, carrying forward a hidden state that summarizes everything seen so far. LSTMs (Long Short-Term Memory networks) are a widely used RNN variant designed to better retain information over longer sequences. RNNs were the dominant sequence-modeling approach before 2017, but their strictly sequential processing makes them slow to train (no parallelism across time steps) and prone to losing long-range information — the exact problems the transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. was designed to solve.
How it works
At each step the network takes the current input and the previous
hidden state and produces a new one, roughly
h_t = tanh(W_h h_prev + W_x x_t + b). The same weights are reused at
every step, so parameter count is independent of sequence length.
Training uses backpropagation through time: the loop is unrolled
into a deep feed-forward graph and gradients flow backward through
every step, so compute and memory scale with sequence length.
An LSTM adds a separate cell state plus input, forget, and output gates that learn what to keep and what to discard. The gated path lets the cell state cross many steps with multiplicative factors near one, which is what makes long dependencies learnable at all.
When it breaks
- Vanishing gradientsVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training.. Repeated multiplication by the same Jacobian shrinks or explodes gradients across many steps. Clipping handles the explosion; backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network. through a gated path only partly handles the vanishing.
- No parallelism over time. Step
tneedsh_{t-1}, so a GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires. sits mostly idle during training — the practical reason transformers won. - Bottlenecked state. The entire prefix is compressed into one fixed-size vector, so early information is squeezed out as the sequence grows.
- Truncation artifacts. Backpropagation through time is usually truncated to fit memory, silently capping the dependency length the model can learn.
See also: TransformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs., AttentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer.
Learn more: Attention & Transformers · Wikipedia: Recurrent neural network
Mentioned in
Lessons where this comes up in context.
Encoder-Decoder Architecture
Encoder-only, decoder-only, and encoder-decoder describe which half of the original transformer a model uses, determining whether it's built for understanding, generation, or both.
LLM (Large Language Model)
An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama.