Gradient Descent
Gradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient.
Gradient descent is the algorithm used to train nearly every model in
modern AI. It repeatedly computes the gradient of the
loss functionLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. with respect to the model's
parameters, then updates the parameters a small step in the opposite
direction (since the gradient points toward increasing loss):
θ ← θ - η · ∇L(θ), where η is the learning rate.
Stochastic gradient descent (SGD) estimates the gradient from a small
random batch of data at a time rather than the full dataset; Adam is
a widely used variant that adapts the effective learning rate per
parameter.
How it works
One training step is: sample a mini-batch, run the forward pass, compute
the loss, call backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network. to get
grad, then apply w := w - lr * grad. Repeat until the loss plateaus.
The practical variants differ in how they turn raw gradients into a step:
- Momentum accumulates an exponentially decayed running average of past gradients, damping oscillation across narrow ravines.
- Adam and AdamW keep running estimates of both the first and second moment of the gradient and divide by the square root of the second, giving each parameter its own effective step size.
- Learning-rate schedules (warmup then cosine or linear decay) are standard for transformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.; a large constant rate rarely converges well.
Batch size trades gradient noise for throughput — smaller batches are noisier but take more steps per epoch.
When it breaks
- Learning rate dominates everything. Too high and the loss spikes
to
NaNwithin a few hundred steps; too low and it decreases so slowly it looks like a modelling problem. It is almost always the first hyperparameter to sweep. - Loss goes to
NaN. Usually an exploding gradient, alog(0)in the loss, or mixed-precision overflow. Gradient clipping and a loss scaler are the standard fixes. - Plateaus are not always minima. Long flat stretches often come from saturated activationsActivation FunctionAn activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function. or a decayed schedule, not from convergence.
- AdamW ≠ Adam + weight decay. Adding L2 to the loss under Adam interacts with the adaptive denominator and regularizes far less than intended.
See also: Loss FunctionLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training., BackpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network.
Learn more: ML Fundamentals · Wikipedia: Gradient descent
Mentioned in
Lessons where this comes up in context.
- AI SecurityAdversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
- Applied & Agentic SystemsHow prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
- Computer VisionConvolutions, pooling, CNNs, transfer learning
- ML FundamentalsSupervised/unsupervised learning, loss functions, gradient descent
- Neural Networks & BackpropFrom Karpathy's micrograd approach — building a tiny neural net and stepping through forward/backward passes
- Optimization & Training DynamicsWhy plain gradient descent isn't what actually trains modern models — Adam's mechanics, why warmup exists, how learning-rate schedules are shaped, and what batch size actually trades off at scale
- Probability & Statistics FoundationsDistributions, Bayes' theorem, and maximum likelihood estimation — the math that loss functions and uncertainty in ML are actually built on
- Reinforcement LearningMDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
Activation Function
An activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function.
Backpropagation
Backpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network.