Data Augmentation
Data augmentation expands a training set by applying label-preserving transformations (crops, flips, color jitter) to existing examples, fighting overfitting for free.
Data augmentation artificially expands a training set by applying label-preserving transformations to existing examples — random crops, flips, rotations, or color jitter for images. A photo of a cat flipped horizontally is still obviously a cat, so this manufactures new, valid training examples at no extra labeling cost, and is standard practice in nearly every CNNCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently. training pipeline as a cheap defense against overfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data..
How it works
Augmentation runs on the fly in the input pipeline, so each epoch sees a
different random perturbation of the same image and the model never
fits a fixed arrangement of pixels. A typical image policy composes
RandomResizedCrop, RandomHorizontalFlip, and color jitter, then
normalization. Two properties matter: the transform must preserve the
label, and the randomness must be resampled per example rather than
fixed once per epoch. Stronger policies mix examples instead of merely
perturbing them — mixup blends two images and their labels linearly,
CutMix pastes a patch of one image into another and mixes the labels
in proportion to area. Functionally this is a
regularizerRegularizationRegularization is any technique that trades some training-data fit for better generalization, fighting overfitting on purpose., and it belongs to the
training split only.
When it breaks
- Applying random transforms to validation or test data silently corrupts your evaluation splitTrain/Validation/Test SplitSplitting data into training, validation, and test sets lets you measure a model's ability to generalize to new data instead of just memorizing what it trained on.; only deterministic resize and normalize belong there.
- An overly aggressive policy destroys signal — a hard crop can remove the object entirely, leaving a confidently mislabeled example.
- The transforms encode assumptions. Horizontal flips are fine for natural photos and wrong for medical scans, where laterality is diagnostic.
- CPU-side augmentation routinely becomes the training bottleneck, leaving the accelerator idle between batches.
See also: OverfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data., CNNCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently.
Learn more: Computer Vision
CLIP (Contrastive Language-Image Pretraining)
CLIP trains an image encoder and a text encoder together with a contrastive loss so that matching image-caption pairs land close together in one shared embedding space.
GAN (Generative Adversarial Network)
A GAN trains a generator and a discriminator against each other, the generator learning to fool the discriminator into mistaking its output for real data.