Train/Validation/Test Split
Splitting data into training, validation, and test sets lets you measure a model's ability to generalize to new data instead of just memorizing what it trained on.
The train/validation/test split divides data into three disjoint sets: the training set a model actually learns from; the validation set, held out to compare architectures and tune hyperparameters during development; and the test set, touched exactly once at the end to report final, unbiased performance. This structure exists specifically to detect overfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data. — performance on data the model has never seen is the only honest measure of whether it actually generalizes.
How it works
A common starting point is roughly 80/10/10, though the right ratio
depends on dataset size — with millions of rows, a few thousand held-out
examples already estimate error tightly, while small datasets often call
for k-fold cross-validation instead, which rotates the validation
fold k times and averages.
The split must happen before any fitting step. Scalers, vocabularies, imputation statistics, and feature selection are all fit on the training split only and then applied to the others; fitting them on the full dataset leaks test information into training.
How you split matters as much as the ratio:
- Time series: split chronologically, never randomly.
- Grouped data (multiple rows per user, patient, or document): split by group.
- Imbalanced classes: stratify so each split has the same class mix.
When it breaks
- Near-duplicates across splits. Augmented copies, re-uploaded images, or boilerplate text appearing in both train and test turn the test score into a memorization score. This is the same contamination problem that affects LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. benchmarks.
- Test-set creep. Every decision made after looking at test results leaks it into the process; after enough iterations it is just a second validation set.
- Random splits on dependent data. Consecutive time steps or multiple rows from the same user land on both sides and inflate the score dramatically.
- A validation set that is not the deployment distribution. It will confidently endorse a model that fails in production.
See also: OverfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data., Bias-Variance TradeoffBias-Variance TradeoffThe bias-variance tradeoff describes the balance between a model too simple to fit the data (high bias) and one that overfits its training data's noise (high variance).
Learn more: ML Fundamentals
Bias-Variance Tradeoff
The bias-variance tradeoff describes the balance between a model too simple to fit the data (high bias) and one that overfits its training data's noise (high variance).
Probability Distribution
A probability distribution assigns a likelihood to each possible value of a random variable — the object every loss function is secretly built to measure the fit of.