Batch/Layer Normalization
Batch and layer normalization rescale a layer's outputs during training to keep values well-behaved as they pass through many stacked layers.
Batch normalization and layer normalization rescale a layer's outputs (to roughly mean 0, variance 1) during training, preventing values from drifting too large or small as they pass through many stacked layers — a direct countermeasure to vanishing/exploding gradientsVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training.. Batch normalization normalizes across a batch of examples; layer normalization normalizes across a single example's features instead, and is the variant used throughout transformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs..
How it works
Both compute y = γ · (x - μ) / sqrt(σ² + ε) + β. What differs is the
axis the statistics are taken over.
- BatchNorm averages over the batch dimension, so for activations of
shape
(N, C, H, W)it produces oneμandσ²per channelC. It also maintains a running mean and variance, used at inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. time when there is no batch to average over — which is why the layer behaves differently intrain()andeval()mode. - LayerNorm averages over the feature dimension of a single example, so it has no batch dependence and no running statistics, and behaves identically in training and inference.
γ and β are learned, letting the network undo the normalization when
that is useful.
When it breaks
- Small or uneven batches. BatchNorm statistics get noisy below roughly 8–16 examples per device, and a batch size of 1 is degenerate. This is why detection and segmentation models often switch to GroupNorm.
- Forgetting
model.eval(). BatchNorm keeps using batch statistics and keeps updating its running averages, so evaluation results shift depending on what else is in the batch. - Train/inference skew. If the deployment data distribution differs from training, the frozen running statistics are simply wrong, and accuracy drops with no change in weights.
- Sequence data. Variable-length padding pollutes batch statistics, one reason transformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. and RNNsRNN (Recurrent Neural Network)An RNN is a neural network for sequences that processes input one step at a time, carrying a hidden state forward — the main predecessor to transformers. use LayerNorm instead.
See also: Vanishing/Exploding GradientsVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training., TransformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.
Learn more: Neural Networks & Backprop
Weight Initialization
Weight initialization sets a neural network's starting parameter values, using schemes like Xavier/He initialization to keep gradients well-scaled from the first training step.
Residual Connection
A residual (skip) connection adds a layer's input back to its output, giving gradients a direct path backward and making very deep networks trainable.