Regularization
Regularization is any technique that trades some training-data fit for better generalization, fighting overfitting on purpose.
Regularization covers techniques that deliberately trade a bit of training-data fit for better generalization, countering overfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data.. Common forms: weight decay (penalizing large parameter values), dropout (randomly zeroing neuron outputs during training), and early stopping (halting training once validation loss stops improving). More (varied) training data has the same effect without being a training-time technique per se.
How it works
Each technique constrains the space of solutions gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. can settle into:
- Weight decay adds a penalty on
||w||**2to the objective, which shrinks every weight slightly on each step and biases the model toward smaller, smoother functions. In modern optimizers useAdamW, where the decay is applied to the weights directly rather than folded into the gradient. - Dropout zeroes each unit's output with probability
pduring training and rescales the survivors, so no single unit can be relied on. It is disabled at inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. time. - Early stopping simply keeps the checkpoint with the best validation loss.
- Data augmentationData AugmentationData augmentation expands a training set by applying label-preserving transformations (crops, flips, color jitter) to existing examples, fighting overfitting for free. expands the effective dataset with label-preserving transforms.
All of them raise training loss on purpose; the payoff is lower validation loss.
When it breaks
- Regularizing an underfit model. If training and validation loss are both high, adding dropout or decay makes both worse. Diagnose the bias-varianceBias-Variance TradeoffThe bias-variance tradeoff describes the balance between a model too simple to fit the data (high bias) and one that overfits its training data's noise (high variance). side first.
- Dropout left on at eval. Forgetting
model.eval()injects noise into predictions and makes evaluation runs non-reproducible. - Decaying the wrong parameters. Applying weight decay to biases and
normalization
γ/βterms is standard-but-wrong; most reference training scripts exclude them. - Stacking too much. Heavy dropout plus strong decay plus aggressive augmentation can jointly prevent the model from fitting at all — tune one at a time.
See also: OverfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data., Neural NetworkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation.
Learn more: ML Fundamentals · Wikipedia: Regularization (mathematics)
Mentioned in
Lessons where this comes up in context.
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- ML FundamentalsSupervised/unsupervised learning, loss functions, gradient descent
- Probability & Statistics FoundationsDistributions, Bayes' theorem, and maximum likelihood estimation — the math that loss functions and uncertainty in ML are actually built on
Overfitting
Overfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data.
Bias-Variance Tradeoff
The bias-variance tradeoff describes the balance between a model too simple to fit the data (high bias) and one that overfits its training data's noise (high variance).