Weight Initialization
Weight initialization sets a neural network's starting parameter values, using schemes like Xavier/He initialization to keep gradients well-scaled from the first training step.
Weight initialization is the choice of starting values for a neural networkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation.'s parameters before training begins. Poor initialization (too large or too small) worsens vanishing/exploding gradientsVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training. from step one. Xavier/Glorot and He initialization are schemes that set the initial random weight scale based on a layer's number of inputs/outputs, specifically to keep activations and gradients at a consistent scale as they pass through the network.
How it works
Both schemes pick a variance that keeps the signal's scale roughly
constant across a layer. For a layer with fan_in inputs and fan_out
outputs:
- Xavier/Glorot uses variance
2 / (fan_in + fan_out), derived assuming a symmetric, roughly linear activation — the right match fortanh. - He uses
2 / fan_in, doubling the variance to compensate for ReLU discarding half its inputs. This is the default for ReLU-family networks (kaiming_normal_in PyTorch).
Biases start at zero. Symmetry matters too: initializing all weights to the same value makes every neuron in a layer compute the same thing and receive the same gradient forever, so they never differentiate — which is why the values must be random, not merely small.
When it breaks
- All-zeros or all-equal weights. The network trains but the layer behaves as if it had one neuron. Symmetry is never broken by gradient descent alone.
- Wrong scheme for the activation. Using Xavier with ReLU halves the signal variance per layer, so deep stacks fade toward zero — exactly the vanishing gradientVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training. pattern, from step zero.
- Framework defaults are not always right. PyTorch's default for
nn.Linearis a uniform bound based onfan_in, not He; deep ReLU models often want an explicit re-init. - Pretrained weights overwritten. Calling a custom init function across the whole model after loading a checkpoint silently destroys the pretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability. you were trying to reuse.
See also: Vanishing/Exploding GradientsVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training., Neural NetworkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation.
Learn more: Neural Networks & Backprop
Mentioned in
Lessons where this comes up in context.
Vanishing/Exploding Gradients
Vanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training.
Batch/Layer Normalization
Batch and layer normalization rescale a layer's outputs during training to keep values well-behaved as they pass through many stacked layers.