Probability Distribution
A probability distribution assigns a likelihood to each possible value of a random variable — the object every loss function is secretly built to measure the fit of.
A probability distribution assigns a likelihood to each possible outcome of a random quantity. It's the mathematical object underneath almost every part of this course that involves uncertainty: a classifier's output, a language model's next-token prediction, and the noise term in a regression model are all distributions, not single numbers.
How it works
Three distributions come up constantly:
- Bernoulli: a single binary outcome (spam or not) with probability
pof "yes." A binary classifier's output is literally thepof a Bernoulli distribution. - Categorical: the multi-class generalization — a probability for each of several discrete outcomes, summing to 1. This is exactly what an LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. outputs at each generation step: a probability for every possible next token.
- Gaussian (normal): the bell curve, defined by a mean and variance, used to model continuous noise and uncertainty — and the assumption behind why mean squared errorLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. is a sensible loss at all.
The expectation of a distribution is its probability-weighted average; variance measures how spread out it is around that average. A loss function is an expectation: cross-entropy is the expected negative log-probability the model assigns to the correct answer, averaged over training examples.
When it breaks
- You never have the true distribution, only a sample. Every statistic computed from data — a sample mean, a benchmark score, a validation accuracy — is an estimate with its own uncertainty, not the ground truth.
- Estimates improve with the square root of sample size, not linearly. Halving the uncertainty on an estimate costs four times the data — the reason a validation set of 100 examples gives a noisy signal and 10,000 gives a usable one.
- A model's output being a probability doesn't mean it's calibrated. Cross-entropy training pushes toward confident predictions, not accurate ones — a model can output 0.95 on inputs it's actually right about only 70% of the time.
See also: Bayes' TheoremBayes' TheoremBayes' theorem relates how likely a hypothesis is given evidence to how likely the evidence is given the hypothesis — the exact mechanism behind spam filters and MAP estimation., MLEMLE (Maximum Likelihood Estimation)MLE chooses model parameters that make the observed training data as probable as possible — and turns out to be exactly what training with cross-entropy or MSE loss already does., Loss FunctionLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training.
Learn more: Probability & Statistics Foundations
Mentioned in
Lessons where this comes up in context.
- Attention & TransformersThe core architecture behind modern AI
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- Probability & Statistics FoundationsDistributions, Bayes' theorem, and maximum likelihood estimation — the math that loss functions and uncertainty in ML are actually built on
- Reinforcement LearningMDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF
Train/Validation/Test Split
Splitting data into training, validation, and test sets lets you measure a model's ability to generalize to new data instead of just memorizing what it trained on.
Bayes' Theorem
Bayes' theorem relates how likely a hypothesis is given evidence to how likely the evidence is given the hypothesis — the exact mechanism behind spam filters and MAP estimation.