Systems

Evaluation & Benchmarks

How models are actually scored — benchmark suites, leaderboards, eval methodology, and why a high benchmark score doesn't guarantee good real-world performance

Every lesson so far has assumed you can tell whether a model is "good." ML Fundamentals covered basic metrics (accuracy, precision, recall) for simple classifiers — but evaluating a modern LLM, which can write code, answer trivia, reason through math, and hold a conversation, is a much harder and more contested problem. This lesson covers how it's actually done, and where it goes wrong.

Picking an evaluation method

Before reaching for a specific benchmark, it helps to ask one question: does the task have a single answer you can check by machine? Almost every evaluation method in this lesson is an answer to some version of that question, and the further "no" you go, the more the evaluation starts to cost.

The left branch is the comfortable one. If the answer is B, or the generated function either passes the unit tests or it doesn't, scoring is exact, free, and reproducible — which is exactly why the best-known benchmarks live there. The right branch is where most of what people actually use LLMs for lives, and where every option is a compromise.

Why isn't accuracy enough to evaluate a model?

When a task does have checkable answers, the metrics from ML Fundamentals apply unchanged: accuracy when the classes are balanced, and precision and recall when they aren't or when the two kinds of mistake cost different amounts. A safety filter that blocks harmful requests, for instance, is judged almost entirely on the precision/recall balance rather than on raw accuracy.

Drag the threshold below and watch precision and recall trade off against each other directly, on a fixed set of scored examples — and watch the same threshold trace out one point on a ROC curve, a plot of true positive rate against false positive rate across every possible threshold at once:

Confusion matrix, 300 scored examples

Predicted +
Predicted −
Actual +
120
30
Actual −
29
121
precision
80.5%
recall
80.0%
F1
80.3%
FPR 19.3%TPR 80.0%AUC 0.88

Drag the threshold down and both true and false positives rise together — the dot slides up the curve. The curve itself doesn't move; only the threshold changes where on it you're standing.

Some tasks sit in between: there's a reference answer, but no single exact string counts as correct — translation and summarization are the classic cases, since two fluent translations of the same sentence can use entirely different words. Reference-based overlap metrics like BLEU and ROUGE score how much word- and phrase-level overlap a generated output shares with the reference text. They're cheap and reproducible, same as exact match, but they reward wording similarity to one specific reference rather than correctness or fluency directly — which is exactly why they've fallen out of favor for anything beyond translation, as better options below became available.

Benchmark suites: standardized exams for models

A benchmark is a fixed, standardized set of test questions used to compare models on a specific capability. A few widely cited ones:

  • MMLU (Massive Multitask Language Understanding): multiple-choice questions spanning dozens of academic subjects, used as a broad knowledge-and-reasoning benchmark.
  • HumanEval: programming problems, scored by whether generated code actually passes test cases — a benchmark for coding ability specifically.
  • GSM8K: grade-school math word problems, testing multi-step arithmetic reasoning of the kind chain-of-thought prompting is meant to help with.

Model releases routinely report scores on suites like these, and leaderboards rank models by aggregate performance across many benchmarks at once. This gives a standardized, reproducible way to compare models — the same appeal standardized tests have for comparing students.

One thing leaderboards rarely show is how precise those scores are. Benchmarks contain a finite number of questions, so a reported score is an estimate from a sample, and like any sample estimate it comes with noise. With a few hundred items, that noise is larger than most of the gaps being argued about.

How big a gap is real?

Take a benchmark with 400 questions and a model scoring 80%. The standard error of a proportion is p(1−p)/n\sqrt{p(1-p)/n}:

  • 0.8×0.2/400=0.0004=0.02\sqrt{0.8 \times 0.2 / 400} = \sqrt{0.0004} = 0.02 → 2 points

So a single score of 80% is really "80% give or take about 4 points" at 95% confidence. Comparing two models on different samples roughly multiplies the uncertainty by 2\sqrt{2}: the gap between them carries a 95% interval of about ±5.5 points.

A headline like "81.2% vs 79.8%" is, on 400 items, indistinguishable from a tie.

Why benchmark scores can mislead

A high score is not the same as being good at the task

Benchmarks measure performance on a fixed, known set of questions — which creates several specific failure modes worth naming.

  • Contamination: if a benchmark's questions (or close variants) appeared in a model's pretraining data, the model can score well by having partially memorized answers rather than by genuinely reasoning — a real and hard-to-fully-rule-out risk for benchmarks built from public web text.
  • Overfitting to the leaderboard: if a lab tunes its models specifically to score well on popular public benchmarks (directly analogous to overfitting to a validation set from ML Fundamentals), benchmark performance can improve without real-world capability improving to match.
  • Narrow coverage: a benchmark measures exactly what it tests — a model that excels at multiple-choice academic questions (MMLU) isn't thereby guaranteed to be good at open-ended writing, following complex multi-step instructions, or knowing when it doesn't know something (see hallucination), since none of those are what that particular benchmark measures.
  • Saturation: as models improve, top scores on a benchmark cluster near 100%, at which point the benchmark stops usefully distinguishing between models — a recurring pattern that drives the field to continually design harder benchmarks.

LLM-as-judge: evaluating open-ended output

Multiple-choice benchmarks are easy to score automatically, but a lot of what makes an LLM useful — writing quality, helpfulness, following nuanced instructions — has no single correct answer to check against. LLM-as-judge evaluation uses a (typically stronger) LLM to score or compare another model's open-ended outputs, following the same intuition behind RLHF: judging the quality of a response is often easier than defining a hard, checkable rule for it. This approach scales far better than paying humans to review every output, but inherits its own risks — a judge model's own biases and blind spots quietly become the evaluation criteria.

Held-out and human evaluation

Beyond automated benchmarks, two other evaluation modes matter in practice:

  • Held-out test sets: the same train/test split discipline from ML Fundamentals applies here too — a genuinely unseen, uncontaminated set of examples is the most trustworthy evidence of real generalization, precisely because it can't have leaked into training data if properly kept private.
  • Human evaluation: humans directly rate or compare model outputs — the same kind of preference data RLHF trains on. It's the closest proxy to real user experience, but is slow, expensive, and doesn't scale to the pace of model iteration the way automated benchmarks do.

Human evaluation is usually collected as pairwise comparisons rather than absolute scores: show a rater two responses to the same prompt and ask which is better. People are far more consistent at ranking two things than at assigning a number to one thing, and the resulting votes can be aggregated into a single rating per model.

MethodCost & speedWhat it misses
Automatic benchmarksNear-free, secondsAnything without a checkable answer; vulnerable to contamination
Model-graded judgingCheap, minutesThe judge's own biases and blind spots become the criteria
Human preferenceExpensive, daysSlow enough that it lags model iteration; raters disagree

In practice, serious model evaluation combines several of these — broad automated benchmarks for quick, cheap signal; held-out and contamination-checked test sets for trustworthy generalization claims; and human or LLM-as-judge evaluation for the open-ended qualities no multiple-choice benchmark can capture.

Recap and what's next

No single number captures "how good" a model is. Benchmark suites give standardized, comparable scores but are vulnerable to contamination, leaderboard-chasing, and narrow coverage; LLM-as-judge and human evaluation cover open-ended quality that benchmarks can't, at a real cost in speed and scale. Treat any single reported score — including ones from this course's own examples — as one data point, not the whole picture. The final lessons look at what gets built on top of a served, evaluated LLM: prompting, retrieval, and multi-step agentic systems.

On this page