Probability & Statistics Foundations
Distributions, Bayes' theorem, and maximum likelihood estimation — the math that loss functions and uncertainty in ML are actually built on
ML Fundamentals introduced loss functionsLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. and gradient descent without explaining why cross-entropy loss is the "right" choice for classification, or what a model's output probabilities actually mean. Both answers come from probability and statistics — not as background trivia, but as the actual mathematical foundation loss functions, uncertainty, and even regularization are built on.
The frame to hold onto: there is some true distribution out in the world that produced your data, you only ever see a finite sample of it, and everything training does is pick the parameters that make that sample look as likely as possible. Every step in the pipeline is a step in that chain.
The arrow looping back is the whole game. You never get to measure how well you did against the true distribution — you only ever get another finite sample, held out, and use it to estimate that gap.
Random variables and distributions
A random variable represents a quantity whose value isn't fixed but follows some pattern of likelihood — a distribution. A probability distribution assigns a likelihood to each possible outcome. Three you'll see constantly in this course:
- Bernoulli distribution: models a single binary outcome (spam / not
spam) with probability
pof "yes." The output of a binary classifier is literally thepof a Bernoulli distribution. - Categorical distribution: the multi-class generalization — a probability for each of several discrete outcomes, summing to 1. This is exactly what a language model outputs at each generation step: a probability for every possible next token (see LLMs).
- Gaussian (normal) distribution: the classic bell curve, defined by a mean and variance, used constantly to model continuous noise and uncertainty, and as the default assumption behind mean squared error loss (below).
Expectation and variance
The expectation (or expected value) of a random variable is its probability-weighted average — the long-run average value if you sampled from the distribution infinitely many times. Variance measures how spread out a distribution is around its expectation. These aren't just abstract definitions: a loss function is an expectation — cross-entropy loss is the expected negative log-probability the model assigns to the correct answer, averaged over your training examples.
But you can't compute that expectation directly, because you don't have the true distribution — you have a finite sample of it. So you compute the sample mean instead and hope it's close. It usually is, and there's a precise reason why.
For a random variable with distribution , the expectation and variance are:
Given independent samples, the sample mean is unbiased — — and its variance shrinks with :
So the standard error — the typical distance between your measured average and the truth — falls as . That square root is the most consequential fact in applied statistics: halving your error bar costs four times the data, and getting 10× more precision costs 100×.
This single expression explains two things at once. It's why a validation set of 100 examples gives a noisy accuracy estimate while 10,000 gives a usable one, and it's why mini-batch gradients work — the same variance reduction applies to gradient estimates, which is covered in ML Fundamentals.
Bayes' theorem
Bayes' theoremBayes' TheoremBayes' theorem relates how likely a hypothesis is given evidence to how likely the evidence is given the hypothesis — the exact mechanism behind spam filters and MAP estimation. relates two conditional probabilities — how likely a hypothesis is given evidence, versus how likely the evidence is given the hypothesis:
P(hypothesis | evidence) = P(evidence | hypothesis) · P(hypothesis)
────────────────────────────────────────
P(evidence)In MLMachine Learning (ML)Machine learning is the practice of writing programs that learn a function from data, via a loss function and an optimizer, instead of following hand-written rules. terms: the prior, P(hypothesis), is what you believed before
seeing data; the likelihood, P(evidence | hypothesis), is how well a
hypothesis explains the data you observed; the posterior,
P(hypothesis | evidence), is your updated belief after seeing it. This
isn't just philosophy — it's the exact mechanism behind spam filters
(is this email spam, given these words?), and the conceptual basis for
the MAPMAP (Maximum a Posteriori Estimation)MAP estimation extends MLE with a prior belief about the parameters themselves, via Bayes' theorem — and is precisely where weight-decay regularization comes from. estimation below.
Bayesian vs. frequentist
These two conditional probabilities look similar but answer different questions, and conflating them is a classic reasoning error: "how likely is a positive test result, given you have the disease" is very different from "how likely is it you have the disease, given a positive test result" — especially when the disease is rare. Bayes' theorem is precisely the tool for converting between the two correctly.
For a hypothesis and observed evidence :
The denominator is just a normalising constant — the total probability of seeing that evidence at all, summed over every hypothesis:
The cleanest way to see what's going on is the odds form. Divide the rule for by the rule for its negation ; the terms cancel:
Evidence multiplies your prior odds by how much better the hypothesis explains it than the alternative does. A test that fires 20× more often on sick people than healthy ones multiplies your odds by 20 — which is a lot, but not enough if your prior odds were 1 in 10,000. Strong evidence cannot rescue a weak prior on its own, and that is the entire content of the base-rate fallacy.
A disease affects 1 in 10,000 people. A test catches 99% of real cases and has a 1% false-positive rate. You test positive. How worried should you be?
Imagine 1,000,000 people tested:
- 100 have the disease. 99 of them test positive.
- 999,900 are healthy. 1% of them — about 9,999 — test positive anyway.
So 10,098 positive results, of which 99 are real:
A "99% accurate" test on a positive result leaves you 99% likely to be fine. Nothing about the test is broken — the false positives simply come from a pool 10,000× larger. Swap disease for fraud, spam, or safety violation and the same arithmetic explains why rare-event classifiers drown in false alarms.
Drag the sliders below to run that same arithmetic on any prior, sensitivity, and false-positive rate — watch how little a positive result moves your belief when the condition is rare, and how much more it moves once the condition is common.
Belief that you have the condition, before vs. after one positive test result:
Out of 1,000,000 people, about 100 actually have the condition. Testing everyone gives 99 true positives and 9,999 false positives — so a positive result means a 0.98% chance of actually having it.
Maximum likelihood estimation: what training actually optimizes
Here's the connection that ties this lesson directly back to
ML Fundamentals: maximum likelihood estimationMLE (Maximum Likelihood Estimation)MLE chooses model parameters that make the observed training data as probable as possible — and turns out to be exactly what training with cross-entropy or MSE loss already does.
(MLE) is the principle of choosing model parameters θ that make the
observed training data as probable as possible under the model.
For classification, this means: choose θ to maximize the probability
the model assigns to the correct label, averaged across all training
examples. Maximizing a probability is equivalent to minimizing its
negative logarithm (logarithm is monotonic, and working with sums of logs
instead of products of probabilities is numerically far more stable) —
and that negative log-probability, averaged over the data, is exactly
cross-entropy loss.
This is not a coincidence or a rough analogy: training a classifier with cross-entropy loss via gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. is maximum likelihood estimation. Every loss function in this course has a probabilistic story behind it — mean squared error is the MLE solution under an assumption that errors are Gaussian-distributed noise.
Assume the training examples are independent. The likelihood of the whole dataset under parameters is the product of the per-example probabilities:
Products of thousands of numbers below 1 underflow to zero in floating point, so take logs — monotonic, so the maximiser is unchanged — and flip the sign to turn maximisation into minimisation:
That right-hand expression is the cross-entropy loss from ML Fundamentals — the negative log-probability assigned to the correct label, averaged over the data. No derivation step was skipped; they are the same object written two ways.
The Gaussian case follows the same route. Assume with . Then , and taking the negative log leaves:
The constant doesn't depend on , so minimising it is minimising mean squared error. MSE is not an arbitrary choice of "squaring feels right" — it is the assumption of Gaussian noise, written as a loss.
MAP estimation: where regularization comes from
Maximum a posteriori (MAP) estimation extends MLE by incorporating a
prior belief about the parameters themselves, via Bayes' theorem: instead
of only maximizing how well θ explains the data (the likelihood), MAP
also weighs how plausible θ was before seeing the data (the prior).
This turns out to be exactly where weight decayRegularizationRegularization is any technique that trades some training-data fit for better generalization, fighting overfitting on purpose. comes from: assuming a Gaussian prior centered at zero for the parameters — a belief that "smaller weights are more plausible than huge ones" — and combining it with the likelihood via Bayes' theorem produces, after the same log-probability trick as above, a training objective that is precisely the ordinary loss function plus an L2 penalty on the weights. Regularization isn't a separate hack bolted onto training — it's what MLE becomes once you add a prior belief about the parameters.
| Maximises | Uses a prior | Training objective | |
|---|---|---|---|
| MLE | Probability of the data given θ | No | Plain loss |
| MAP | Probability of θ given the data | Yes | Loss + regularization penalty |
The choice of prior picks the penalty: a Gaussian prior gives the squared (L2) penalty of weight decay, while a Laplace prior gives an absolute-value (L1) penalty, which drives weights to exactly zero and produces sparse models.
Recap and what's next
Probability distributionsProbability DistributionA probability distribution assigns a likelihood to each possible value of a random variable — the object every loss function is secretly built to measure the fit of. are the objects loss functions are built to measure the fit of; maximum likelihood estimation is the principle that justifies cross-entropy and MSE loss as "correct" rather than arbitrary choices; and MAP estimation shows that regularization is a special case of the same framework, not a separate trick. Keep this lens handy — whenever a later lesson introduces a loss function or a training objective, it has a probabilistic story like the ones here underneath it. The next lesson returns to the mechanics of training itself: how gradients are actually computed through a network, from scratch.
Optimization & Training Dynamics
Why plain gradient descent isn't what actually trains modern models — Adam's mechanics, why warmup exists, how learning-rate schedules are shaped, and what batch size actually trades off at scale
Neural Networks & Backprop
From Karpathy's micrograd approach — building a tiny neural net and stepping through forward/backward passes