Language Models Are Few-Shot Learners
Brown et al., 2020 (the "GPT-3" paper) — showed that scaling a language model to 175 billion parameters let it perform new tasks from just a few examples in the prompt, with no gradient updates at all.
"Language Models Are Few-Shot Learners" (Tom B. Brown, Benjamin Mann, Nick Ryder, and colleagues at OpenAI, 2020) — the GPT-3 paper — demonstrated that a large enough language model could learn a new task from a handful of examples placed directly in its prompt, without a single weight being updated.
What problem it solved
Before this paper, adapting a pretrained language model to a specific task normally meant fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions.: collecting a labeled dataset for that task and updating the model's weights on it. That approach works, but it requires a labeled dataset and a training run per task — impractical for the kind of general-purpose system that should handle whatever a user asks next, and a real barrier to anyone without the infrastructure to run training jobs. The open question: does a large enough pretrained model already contain enough general capability that it can pick up a new task just by being shown a few examples, without any further training at all?
The key idea
Scale the model far beyond what had been tried before — to 175 billion parameters, roughly 10× larger than the largest prior language models — and test it purely through in-context learning: at inference time, place a short description of the task and a handful of example input-output pairs directly in the prompt, then ask the model to continue the pattern on a new input. No gradients, no weight updates, no separate training run — the "learning" happens entirely within the forward pass, just from the examples sitting in the context window. The paper tested this across three settings of decreasing information — few-shot (several examples), one-shot (a single example), and zero-shot (a task description with no examples at all) — and found performance improved smoothly and substantially with scale across all three.
GPT-2 (2019), the previous generation, had 1.5 billion parameters. GPT-3 has 175 billion — roughly a 117× increase in a single generation, trained on a correspondingly larger slice of internet text. That jump is a concrete instance of the pattern the scaling lawsScaling LawsScaling laws are empirical relationships between a model's loss and its parameter count, dataset size, and compute budget, used to plan how large to train a new model. work was starting to formalize around the same time: capability climbing smoothly and predictably as model size, data, and compute all scale together.
Why it mattered
GPT-3 was the first model to make in-context, few-shot learning a practical way to use a language model, rather than a research curiosity — you could point it at a new task by writing a good prompt instead of collecting a labeled dataset and running a training job. That's the capability prompt engineering is built on top of, and it reframed what "using" a language model even meant: not training a specialized model per task, but steering one general model with instructions and examples. GPT-3 itself was still difficult to control reliably and prone to unhelpful or off-target completions — the gap that InstructGPT was built two years later specifically to close.
Authors: Tom B. Brown, Benjamin Mann, Nick Ryder, and colleagues (OpenAI)
Read the paper: arXiv:2005.14165
Learn more: History & Landscape · Prompt Engineering · LLMs
Attention Is All You Need
Vaswani et al., 2017 — introduced the transformer, dropping recurrence and convolution entirely in favor of self-attention. The architecture behind essentially every modern LLM.
LoRA: Low-Rank Adaptation of Large Language Models
Hu et al., 2021 — made fine-tuning huge models affordable by freezing the original weights and training a tiny pair of low-rank matrices alongside them instead.