Benchmark
A benchmark is a fixed, standardized set of test questions used to compare models on a specific capability — reproducible, but vulnerable to contamination and saturation.
A benchmark is a fixed, standardized set of test questions used to compare models on a specific capability. Widely cited examples include MMLU (broad academic knowledge and reasoning), HumanEval (programming problems scored by whether generated code passes test cases), and GSM8K (grade-school math word problems). Leaderboards rank models by aggregate performance across many benchmarks at once.
How it works
Benchmarks work best when a task has a single, machine-checkable answer: exact match on a multiple-choice question, or code that either passes its unit tests or doesn't. That makes scoring exact, free, and reproducible — the same appeal standardized tests have for comparing students. A reported score is nonetheless an estimate from a finite sample of questions, and like any sample estimate it carries statistical uncertainty: with only a few hundred items, that noise is often larger than the gap between two models being compared.
When it breaks
- Contamination. If a benchmark's questions (or close variants) appeared in a model's pretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability. data, the model can score well by partial memorization rather than genuine reasoning — a real risk for any benchmark built from public web text, and hard to fully rule out from outside a lab.
- Overfitting to the leaderboard. If a lab tunes specifically to score well on popular public benchmarks, performance can improve on paper without real-world capability improving to match — the same mechanism as overfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data. to a validation set, distributed across an entire industry (Goodhart's law).
- Narrow coverage. A benchmark measures exactly what it tests; a model that excels at multiple-choice academic questions isn't thereby good at open-ended writing or knowing when it doesn't know something.
- Saturation. As models improve, top scores cluster near 100%, at which point the benchmark stops usefully distinguishing between models.
See also: LLM-as-JudgeLLM-as-JudgeLLM-as-judge uses a typically stronger LLM to score or compare another model's open-ended output, scaling far better than human review at the cost of inheriting the judge's own biases., OverfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data.
Learn more: Evaluation & Benchmarks
Mentioned in
Lessons where this comes up in context.
- Applied & Agentic SystemsHow prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
- Evaluation & BenchmarksHow models are actually scored — benchmark suites, leaderboards, eval methodology, and why a high benchmark score doesn't guarantee good real-world performance
- Probability & Statistics FoundationsDistributions, Bayes' theorem, and maximum likelihood estimation — the math that loss functions and uncertainty in ML are actually built on
MCP (Model Context Protocol)
MCP is an open standard, introduced by Anthropic, for how applications expose tools and data to LLMs, so a tool built once can be reused across different LLM apps.
LLM-as-Judge
LLM-as-judge uses a typically stronger LLM to score or compare another model's open-ended output, scaling far better than human review at the cost of inheriting the judge's own biases.