Transformers & LLMs

Language Models Are Few-Shot Learners

Brown et al., 2020 (the "GPT-3" paper) — showed that scaling a language model to 175 billion parameters let it perform new tasks from just a few examples in the prompt, with no gradient updates at all.

"Language Models Are Few-Shot Learners" (Tom B. Brown, Benjamin Mann, Nick Ryder, and colleagues at OpenAI, 2020) — the GPT-3 paper — demonstrated that a large enough language model could learn a new task from a handful of examples placed directly in its prompt, without a single weight being updated.

What problem it solved

Before this paper, adapting a pretrained language model to a specific task normally meant fine-tuning: collecting a labeled dataset for that task and updating the model's weights on it. That approach works, but it requires a labeled dataset and a training run per task — impractical for the kind of general-purpose system that should handle whatever a user asks next, and a real barrier to anyone without the infrastructure to run training jobs. The open question: does a large enough pretrained model already contain enough general capability that it can pick up a new task just by being shown a few examples, without any further training at all?

The key idea

Scale the model far beyond what had been tried before — to 175 billion parameters, roughly 10× larger than the largest prior language models — and test it purely through in-context learning: at inference time, place a short description of the task and a handful of example input-output pairs directly in the prompt, then ask the model to continue the pattern on a new input. No gradients, no weight updates, no separate training run — the "learning" happens entirely within the forward pass, just from the examples sitting in the context window. The paper tested this across three settings of decreasing information — few-shot (several examples), one-shot (a single example), and zero-shot (a task description with no examples at all) — and found performance improved smoothly and substantially with scale across all three.

How much bigger than what came before

GPT-2 (2019), the previous generation, had 1.5 billion parameters. GPT-3 has 175 billion — roughly a 117× increase in a single generation, trained on a correspondingly larger slice of internet text. That jump is a concrete instance of the pattern the scaling laws work was starting to formalize around the same time: capability climbing smoothly and predictably as model size, data, and compute all scale together.

Why it mattered

GPT-3 was the first model to make in-context, few-shot learning a practical way to use a language model, rather than a research curiosity — you could point it at a new task by writing a good prompt instead of collecting a labeled dataset and running a training job. That's the capability prompt engineering is built on top of, and it reframed what "using" a language model even meant: not training a specialized model per task, but steering one general model with instructions and examples. GPT-3 itself was still difficult to control reliably and prone to unhelpful or off-target completions — the gap that InstructGPT was built two years later specifically to close.

Authors: Tom B. Brown, Benjamin Mann, Nick Ryder, and colleagues (OpenAI)

Read the paper: arXiv:2005.14165

Learn more: History & Landscape · Prompt Engineering · LLMs

On this page