MLE (Maximum Likelihood Estimation)
MLE chooses model parameters that make the observed training data as probable as possible — and turns out to be exactly what training with cross-entropy or MSE loss already does.
MLE (Maximum Likelihood Estimation) is the principle of choosing
model parameters θ that make the observed training data as probable as
possible under the model. It's the answer to a question every earlier
lesson leaves implicit: why is cross-entropyLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training.
the "right" loss for classification, rather than an arbitrary choice?
How it works
For classification, MLE means choosing θ to maximize the probability
the model assigns to the correct label, averaged across training
examples. Maximizing a probability is equivalent to minimizing its
negative logarithm — logs turn products of many small probabilities into
sums, which is numerically stable and doesn't change which θ wins. That
negative log-probability, averaged over the data, is cross-entropy
loss. Training a classifier with cross-entropy via
gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. is not analogous to MLE —
it is MLE, the same object derived two ways.
The same trick explains mean squared error: assuming prediction errors are Gaussian-distributed noise, the negative log-likelihood reduces exactly to MSE. Every loss function in this course has a probabilistic story behind it, not just an intuitive one.
When it breaks
- MLE has no opinion about the parameters themselves. It only asks
how well
θexplains the data, which is exactly what lets it overfit — a model can maximize likelihood on the training set by memorizing noise. MAP estimationMAP (Maximum a Posteriori Estimation)MAP estimation extends MLE with a prior belief about the parameters themselves, via Bayes' theorem — and is precisely where weight-decay regularization comes from. fixes this by adding a prior. - It assumes the model family is capable of representing the truth. If the true relationship isn't Gaussian noise around a linear function, fitting MSE via MLE is optimizing the wrong objective, however well the optimization itself converges.
- A well-fit likelihood doesn't imply calibrated probabilities. MLE optimizes an average over the training set; it doesn't guarantee the model's confidence matches its accuracy on any individual input.
See also: MAPMAP (Maximum a Posteriori Estimation)MAP estimation extends MLE with a prior belief about the parameters themselves, via Bayes' theorem — and is precisely where weight-decay regularization comes from., Bayes' TheoremBayes' TheoremBayes' theorem relates how likely a hypothesis is given evidence to how likely the evidence is given the hypothesis — the exact mechanism behind spam filters and MAP estimation., Loss FunctionLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training., Probability DistributionProbability DistributionA probability distribution assigns a likelihood to each possible value of a random variable — the object every loss function is secretly built to measure the fit of.
Learn more: Probability & Statistics Foundations
Mentioned in
Lessons where this comes up in context.
- Applied & Agentic SystemsHow prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
- Probability & Statistics FoundationsDistributions, Bayes' theorem, and maximum likelihood estimation — the math that loss functions and uncertainty in ML are actually built on
Bayes' Theorem
Bayes' theorem relates how likely a hypothesis is given evidence to how likely the evidence is given the hypothesis — the exact mechanism behind spam filters and MAP estimation.
MAP (Maximum a Posteriori Estimation)
MAP estimation extends MLE with a prior belief about the parameters themselves, via Bayes' theorem — and is precisely where weight-decay regularization comes from.