Overfitting
Overfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data.
Overfitting happens when a model learns its training data — including its noise and quirks — so closely that performance on new, unseen data suffers. The telltale sign: training loss keeps dropping while validation loss stalls or rises. The opposite failure, underfitting, is when a model is too simple to capture even the patterns present in the training data. RegularizationRegularizationRegularization is any technique that trades some training-data fit for better generalization, fighting overfitting on purpose. techniques (dropout, weight decay, early stopping) exist specifically to fight overfitting.
How it works
A model has enough capacity to represent many functions that fit the training points equally well. Without a preference among them, gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. will happily pick one that also encodes the sampling noise, because reducing training loss is the only thing it is asked to do.
You detect it by tracking training and validation loss on a train/validation/test splitTrain/Validation/Test SplitSplitting data into training, validation, and test sets lets you measure a model's ability to generalize to new data instead of just memorizing what it trained on. on the same axes:
- Both high and flat → underfitting; add capacity or train longer.
- Both falling together → healthy; keep going.
- Training falling, validation flat or rising → overfitting; the epoch where validation bottoms out is the early-stopping point.
The gap widens as capacity grows relative to dataset size, which is the bias-variance tradeoffBias-Variance TradeoffThe bias-variance tradeoff describes the balance between a model too simple to fit the data (high bias) and one that overfits its training data's noise (high variance). in practice.
When it breaks
- Validation set overfitting. Tuning hyperparameters against the same validation set hundreds of times overfits it too, so the validation score stops predicting test performance.
- Leakage disguises it. If duplicates or target-derived features cross the split, validation loss tracks training loss and the model looks perfectly healthy until deployment.
- Small validation sets. Below a few hundred examples the curve is noisy enough that early stopping triggers on randomness.
- Noisy labels. With a meaningful fraction of wrong labels, enough capacity will memorize the errors specifically; label cleaning beats any amount of regularization.
See also: RegularizationRegularizationRegularization is any technique that trades some training-data fit for better generalization, fighting overfitting on purpose., Loss FunctionLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training.
Learn more: ML Fundamentals · Wikipedia: Overfitting
Mentioned in
Lessons where this comes up in context.
- Computer VisionConvolutions, pooling, CNNs, transfer learning
- Evaluation & BenchmarksHow models are actually scored — benchmark suites, leaderboards, eval methodology, and why a high benchmark score doesn't guarantee good real-world performance
- ML FundamentalsSupervised/unsupervised learning, loss functions, gradient descent