Scaling Laws
Scaling laws are empirical relationships between a model's loss and its parameter count, dataset size, and compute budget, used to plan how large to train a new model.
Scaling laws empirically relate a language model's final loss to three quantities: parameter count, dataset size (tokens trained on), and compute budget. DeepMind's 2022 "Chinchilla" paper found that many earlier large models were under-trained relative to their size — for a fixed compute budget, a smaller model trained on more data often beats a larger model trained on less. This directly shapes how labs allocate compute between making a model bigger versus training it on more data.
How it works
Should I train a bigger model or feed it more data?
The empirical finding is that test lossLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training.
follows a power law in each of parameter count N, training tokens
D, and compute C — a straight line on a log-log plot. Because
C ≈ 6 * N * D for transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. training,
fixing a compute budget fixes a trade-off curve between model size and
data, and that curve has a minimum. Chinchilla's fit placed the optimum
near roughly 20 training tokens per parameter, far more data per
parameter than earlier models used.
In practice labs train a series of small models, fit the exponents, and extrapolate to pick the configuration for the single expensive run. Hyperparameters such as learning rate and batch size are fit along the same curves.
When it breaks
- Compute-optimal is not deployment-optimal. The Chinchilla point minimizes training loss per FLOP and ignores inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. cost. For a model serving many requests, a smaller model trained well past the optimum is cheaper overall — which is why open models are routinely over-trained.
- Loss is not capability. Smooth loss curves coexist with abrupt changes on downstream benchmarks, so a predicted loss does not tell you what the model will be able to do.
- The fits assume clean, non-repeated data. Returns fall off once a corpus is repeated many times, and the frontier is increasingly data-limited rather than compute-limited.
See also: LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama., PretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability.
Learn more: LLMs · History & Landscape · Paper: Training Compute-Optimal Large Language Models
Mentioned in
Lessons where this comes up in context.
- Attention & TransformersThe core architecture behind modern AI
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
DPO (Direct Preference Optimization)
DPO aligns an LLM to human preferences directly from ranked response pairs, without the separate reward model and reinforcement learning loop RLHF requires.
Mixture of Experts
A mixture-of-experts layer routes each input to only a handful of specialized subnetworks, growing a model's total parameter count without a proportional rise in compute per input.