Evaluation & Benchmarks
How models are actually scored — benchmark suites, leaderboards, eval methodology, and why a high benchmark score doesn't guarantee good real-world performance
Every lesson so far has assumed you can tell whether a model is "good." ML Fundamentals covered basic metrics (accuracy, precision, recall) for simple classifiers — but evaluating a modern LLM, which can write code, answer trivia, reason through math, and hold a conversation, is a much harder and more contested problem. This lesson covers how it's actually done, and where it goes wrong.
Picking an evaluation method
Before reaching for a specific benchmarkBenchmarkA benchmark is a fixed, standardized set of test questions used to compare models on a specific capability — reproducible, but vulnerable to contamination and saturation., it helps to ask one question: does the task have a single answer you can check by machine? Almost every evaluation method in this lesson is an answer to some version of that question, and the further "no" you go, the more the evaluation starts to cost.
The left branch is the comfortable one. If the answer is B, or the
generated function either passes the unit tests or it doesn't, scoring is
exact, free, and reproducible — which is exactly why the best-known
benchmarks live there. The right branch is where most of what people
actually use LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. for lives, and where every option is a compromise.
Why isn't accuracy enough to evaluate a model?
When a task does have checkable answers, the metrics from ML Fundamentals apply unchanged: accuracy when the classes are balanced, and precision and recall when they aren't or when the two kinds of mistake cost different amounts. A safety filter that blocks harmful requests, for instance, is judged almost entirely on the precision/recall balance rather than on raw accuracy.
Count the four outcomes of a binary decision: true positives , false positives , true negatives , and false negatives . Then:
Precision asks of everything I flagged, how much should I have flagged — its denominator counts the model's positive predictions. Recall asks of everything I should have flagged, how much did I catch — its denominator counts the true positives that exist in the data. Neither one references , which is why both stay informative when negatives vastly outnumber positives.
The score is their harmonic mean:
The harmonic mean, unlike the arithmetic mean, is dragged down hard by the smaller of the two: a model with precision and recall has an arithmetic mean of but an of about . That is the point — refuses to rewardRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one. a model that wins one metric by abandoning the other.
Which to optimise is a product decision, not a statistical one. Weighting recall times as important as precision generalises to:
with favouring recall (missing a case is worse) and favouring precision (a false alarm is worse).
Drag the threshold below and watch precision and recall trade off against each other directly, on a fixed set of scored examples — and watch the same threshold trace out one point on a ROC curve, a plot of true positive rate against false positive rate across every possible threshold at once:
Confusion matrix, 300 scored examples
80.5%recall
80.0%F1
80.3%
Drag the threshold down and both true and false positives rise together — the dot slides up the curve. The curve itself doesn't move; only the threshold changes where on it you're standing.
Some tasks sit in between: there's a reference answer, but no single exact string counts as correct — translation and summarization are the classic cases, since two fluent translations of the same sentence can use entirely different words. Reference-based overlap metrics like BLEU and ROUGE score how much word- and phrase-level overlap a generated output shares with the reference text. They're cheap and reproducible, same as exact match, but they reward wording similarity to one specific reference rather than correctness or fluency directly — which is exactly why they've fallen out of favor for anything beyond translation, as better options below became available.
Benchmark suites: standardized exams for models
A benchmark is a fixed, standardized set of test questions used to compare models on a specific capability. A few widely cited ones:
- MMLU (Massive Multitask Language Understanding): multiple-choice questions spanning dozens of academic subjects, used as a broad knowledge-and-reasoning benchmark.
- HumanEval: programming problems, scored by whether generated code actually passes test cases — a benchmark for coding ability specifically.
- GSM8K: grade-school math word problems, testing multi-step arithmetic reasoning of the kind chain-of-thought prompting is meant to help with.
Model releases routinely report scores on suites like these, and leaderboards rank models by aggregate performance across many benchmarks at once. This gives a standardized, reproducible way to compare models — the same appeal standardized tests have for comparing students.
One thing leaderboards rarely show is how precise those scores are. Benchmarks contain a finite number of questions, so a reported score is an estimate from a sample, and like any sample estimate it comes with noise. With a few hundred items, that noise is larger than most of the gaps being argued about.
Take a benchmark with 400 questions and a model scoring 80%. The standard error of a proportion is :
- → 2 points
So a single score of 80% is really "80% give or take about 4 points" at 95% confidence. Comparing two models on different samples roughly multiplies the uncertainty by : the gap between them carries a 95% interval of about ±5.5 points.
A headline like "81.2% vs 79.8%" is, on 400 items, indistinguishable from a tie.
Treat each question as a Bernoulli trial the model passes with probability . The observed score on items has standard error:
and an approximate 95% confidence interval of . The in the denominator is the expensive part: halving the uncertainty requires four times the questions.
To resolve a genuine one-point difference — a margin of around — you need roughly
items, an order of magnitude more than many popular benchmarks contain.
There is one large caveat in the useful direction. The estimate above assumes the two models were measured independently, but benchmark comparisons are paired: both models answer the same questions. Because model scores are strongly correlated across items — hard questions are hard for everyone — a paired test (McNemar's test, or a bootstrap over items) looks only at the questions where the two models disagree and is substantially more sensitive than comparing two independent intervals.
Why benchmark scores can mislead
A high score is not the same as being good at the task
Benchmarks measure performance on a fixed, known set of questions — which creates several specific failure modes worth naming.
- Contamination: if a benchmark's questions (or close variants) appeared in a model's pretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability. data, the model can score well by having partially memorized answers rather than by genuinely reasoning — a real and hard-to-fully-rule-out risk for benchmarks built from public web text.
- Overfitting to the leaderboard: if a lab tunes its models specifically to score well on popular public benchmarks (directly analogous to overfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data. to a validation set from ML Fundamentals), benchmark performance can improve without real-world capability improving to match.
- Narrow coverage: a benchmark measures exactly what it tests — a model that excels at multiple-choice academic questions (MMLU) isn't thereby guaranteed to be good at open-ended writing, following complex multi-step instructions, or knowing when it doesn't know something (see hallucinationHallucinationA hallucination is a confidently stated, fluent LLM output that is factually wrong, a direct consequence of models being trained to produce plausible text rather than verified facts.), since none of those are what that particular benchmark measures.
- Saturation: as models improve, top scores on a benchmark cluster near 100%, at which point the benchmark stops usefully distinguishing between models — a recurring pattern that drives the field to continually design harder benchmarks.
LLM-as-judge: evaluating open-ended output
Multiple-choice benchmarks are easy to score automatically, but a lot of what makes an LLM useful — writing quality, helpfulness, following nuanced instructions — has no single correct answer to check against. LLM-as-judge evaluation uses a (typically stronger) LLM to score or compare another model's open-ended outputs, following the same intuition behind RLHF: judging the quality of a response is often easier than defining a hard, checkable rule for it. This approach scales far better than paying humans to review every output, but inherits its own risks — a judge model's own biases and blind spots quietly become the evaluation criteria.
Held-out and human evaluation
Beyond automated benchmarks, two other evaluation modes matter in practice:
- Held-out test sets: the same train/test splitTrain/Validation/Test SplitSplitting data into training, validation, and test sets lets you measure a model's ability to generalize to new data instead of just memorizing what it trained on. discipline from ML Fundamentals applies here too — a genuinely unseen, uncontaminated set of examples is the most trustworthy evidence of real generalization, precisely because it can't have leaked into training data if properly kept private.
- Human evaluation: humans directly rate or compare model outputs — the same kind of preference data RLHF trains on. It's the closest proxy to real user experience, but is slow, expensive, and doesn't scale to the pace of model iteration the way automated benchmarks do.
Human evaluation is usually collected as pairwise comparisons rather than absolute scores: show a rater two responses to the same prompt and ask which is better. People are far more consistent at ranking two things than at assigning a number to one thing, and the resulting votes can be aggregated into a single rating per model.
The standard aggregation is the Bradley–Terry model, usually reported in Elo units borrowed from chess. Each model gets a latent strength , and the probability that beats is assumed to be:
The constant is pure convention: it sets a 400-point gap to mean about a 10-to-1 win rate, and a 100-point gap to mean roughly 64%.
Given a pile of recorded comparisons, the ratings are whatever values make the observed outcomes most likely — a logistic regression whose features are which two models played. Live leaderboards often use the online Elo update instead, adjusting after each vote:
where is the actual result and controls how fast ratings move. This is convenient but order-dependent; the batch maximum-likelihood fit is more stable when all the votes are already in.
Two assumptions are doing real work here. First, that a single number per model suffices — Bradley–Terry cannot represent non-transitive preferences, where beats , beats , and beats . Second, that ratings are comparable across the prompt mix voters happened to submit; change the distribution of questions and the ranking can change with it.
| Method | Cost & speed | What it misses |
|---|---|---|
| Automatic benchmarks | Near-free, seconds | Anything without a checkable answer; vulnerable to contamination |
| Model-graded judging | Cheap, minutes | The judge's own biases and blind spots become the criteria |
| Human preference | Expensive, days | Slow enough that it lags model iteration; raters disagree |
In practice, serious model evaluation combines several of these — broad automated benchmarks for quick, cheap signal; held-out and contamination-checked test sets for trustworthy generalization claims; and human or LLM-as-judge evaluation for the open-ended qualities no multiple-choice benchmark can capture.
Recap and what's next
No single number captures "how good" a model is. Benchmark suites give standardized, comparable scores but are vulnerable to contamination, leaderboard-chasing, and narrow coverage; LLM-as-judge and human evaluation cover open-ended quality that benchmarks can't, at a real cost in speed and scale. Treat any single reported score — including ones from this course's own examples — as one data point, not the whole picture. The final lessons look at what gets built on top of a served, evaluated LLM: prompting, retrieval, and multi-step agentic systems.