Probability & Statistics

Probability Distribution

A probability distribution assigns a likelihood to each possible value of a random variable — the object every loss function is secretly built to measure the fit of.

A probability distribution assigns a likelihood to each possible outcome of a random quantity. It's the mathematical object underneath almost every part of this course that involves uncertainty: a classifier's output, a language model's next-token prediction, and the noise term in a regression model are all distributions, not single numbers.

How it works

Three distributions come up constantly:

  • Bernoulli: a single binary outcome (spam or not) with probability p of "yes." A binary classifier's output is literally the p of a Bernoulli distribution.
  • Categorical: the multi-class generalization — a probability for each of several discrete outcomes, summing to 1. This is exactly what an LLM outputs at each generation step: a probability for every possible next token.
  • Gaussian (normal): the bell curve, defined by a mean and variance, used to model continuous noise and uncertainty — and the assumption behind why mean squared error is a sensible loss at all.

The expectation of a distribution is its probability-weighted average; variance measures how spread out it is around that average. A loss function is an expectation: cross-entropy is the expected negative log-probability the model assigns to the correct answer, averaged over training examples.

When it breaks

  • You never have the true distribution, only a sample. Every statistic computed from data — a sample mean, a benchmark score, a validation accuracy — is an estimate with its own uncertainty, not the ground truth.
  • Estimates improve with the square root of sample size, not linearly. Halving the uncertainty on an estimate costs four times the data — the reason a validation set of 100 examples gives a noisy signal and 10,000 gives a usable one.
  • A model's output being a probability doesn't mean it's calibrated. Cross-entropy training pushes toward confident predictions, not accurate ones — a model can output 0.95 on inputs it's actually right about only 70% of the time.

See also: Bayes' Theorem, MLE, Loss Function

Learn more: Probability & Statistics Foundations

Mentioned in

Lessons where this comes up in context.

On this page