Attention Is All You Need
Vaswani et al., 2017 — introduced the transformer, dropping recurrence and convolution entirely in favor of self-attention. The architecture behind essentially every modern LLM.
"Attention Is All You Need" (Ashish Vaswani, Noam Shazeer, Niki Parmar, and colleagues, 2017) introduced the transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. — arguably the single most consequential architecture paper in the current era of AI, and the direct ancestor of every major LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. in use today.
What problem it solved
The dominant approach to sequence modeling before this paper — RNNsRNN (Recurrent Neural Network)An RNN is a neural network for sequences that processes input one step at a time, carrying a hidden state forward — the main predecessor to transformers. and their gated variants (LSTMs, GRUs) — processed a sequence one token at a time, each step depending on the previous one's output. That recurrence was a genuine bottleneck in two ways: it made training slow, since positions couldn't be computed in parallel, and it made it hard for information from early in a long sequence to still influence predictions much later, since it had to survive being passed through every intermediate step. Prior work had already shown that adding an attention mechanism on top of a recurrent model improved results — the open question this paper asked was more radical: what if you removed the recurrence entirely and kept only attention?
The key idea
Replace recurrence with self-attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer.: a mechanism where every position in a sequence directly computes how much to "attend to" every other position, in a single parallel operation, instead of information having to flow step by step through a chain. Since removing recurrence also removes any built-in notion of word order, the model adds explicit positional encodings to each token so it still knows where in the sequence each one falls. Stack several layers of this — self-attention followed by a small feed-forward network, repeated — and you have the transformer: no recurrence, no convolution, and (as the title insists) attention doing essentially all of the representational work.
An RNN over a sequence of length requires sequential steps that can't be parallelized — step can't start until step finishes. A transformer's self-attention computes all pairwise interactions in one parallel matrix operation: sequential steps instead of . On modern GPUs, built for exactly this kind of parallel computation, that difference turned training runs that would take weeks into ones that take days — a large part of why transformers scaled as far as they have.
Why it mattered
Every major LLM built since — GPT, Claude, Gemini, LLaMA, and the rest — is a descendant of the architecture this paper introduced, typically using only the decoder half of the original encoder-decoder design. It also turned out the idea wasn't language-specific at all: the same architecture, with only the input representation changed, became the basis for ViT in vision and for models across audio, biology, and other domains entirely. Few architecture papers have had their central idea reused this widely, this directly, this many years later.
Authors: Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin (Google)
Read the paper: arXiv:1706.03762
See also: Ashish VaswaniAshish VaswaniLead author of "Attention Is All You Need" (2017), the paper that introduced the transformer architecture behind every modern LLM. · Noam ShazeerNoam ShazeerCo-authored "Attention Is All You Need" and pioneered sparse mixture-of-experts models, then left Google to found Character.AI before returning to lead work on Gemini.
Learn more: TransformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. · Attention & Transformers
Mentioned in
Lessons where this comes up in context.
Generative Adversarial Networks
Goodfellow et al., 2014 — proposed training two networks against each other, a generator and a discriminator, as a way to learn to generate realistic data. Dominated image generation for most of the following decade.
Language Models Are Few-Shot Learners
Brown et al., 2020 (the "GPT-3" paper) — showed that scaling a language model to 175 billion parameters let it perform new tasks from just a few examples in the prompt, with no gradient updates at all.