Bias-Variance Tradeoff
The bias-variance tradeoff describes the balance between a model too simple to fit the data (high bias) and one that overfits its training data's noise (high variance).
The bias-variance tradeoff formalizes the balance behind overfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data. and underfitting: a model too simple has high bias (systematically wrong, underfitting); a model too flexible has high variance (fits the specific quirks of whatever training data it saw, overfitting). The practical goal in machine learningMachine Learning (ML)Machine learning is the practice of writing programs that learn a function from data, via a loss function and an optimizer, instead of following hand-written rules. is the sweet spot between the two, not the maximum-capacity model.
How it works
Expected prediction error on a new point decomposes into three parts: bias² (how far the average model is from the truth), variance (how much the fitted model changes if you resample the training set), and irreducible noise (the floor set by the data itself, which no model can beat).
Capacity moves bias and variance in opposite directions. A degree-1 polynomial fit to a curved relationship has high bias and near-zero variance; a degree-20 fit traces the noise exactly, giving low bias and high variance. You diagnose which side you are on by comparing training and validation error on a train/validation/test splitTrain/Validation/Test SplitSplitting data into training, validation, and test sets lets you measure a model's ability to generalize to new data instead of just memorizing what it trained on.: both high means bias, a wide gap means variance. RegularizationRegularizationRegularization is any technique that trades some training-data fit for better generalization, fighting overfitting on purpose. buys variance reduction at the cost of a little bias.
When it breaks
- The clean U-curve does not always hold. Very large overparameterized networks often improve again past the interpolation point ("double descent"), so "more parameters means more variance" is not a reliable rule for deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones..
- Misreading the gap. A small train/validation gap with high error on both is an underfitting problem; adding dropout or weight decay makes it worse, not better.
- Leaked or non-representative validation data. The decomposition only guides you if validation error actually estimates generalization; duplicates across splits hide variance entirely.
- Distribution shift. Neither term accounts for test data drawn from a different distribution than training data.
See also: OverfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data., RegularizationRegularizationRegularization is any technique that trades some training-data fit for better generalization, fighting overfitting on purpose.
Learn more: ML Fundamentals · Wikipedia: Bias–variance tradeoff
Mentioned in
Lessons where this comes up in context.
Regularization
Regularization is any technique that trades some training-data fit for better generalization, fighting overfitting on purpose.
Train/Validation/Test Split
Splitting data into training, validation, and test sets lets you measure a model's ability to generalize to new data instead of just memorizing what it trained on.