ML Fundamentals
Supervised/unsupervised learning, loss functions, gradient descent
Every idea in the rest of this course — neural networksNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation., attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer., LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. — is a variation on one core recipe. Learn this recipe well and the rest is mostly "same idea, fancier function."
The recipe: model, loss, optimize
Machine learningMachine Learning (ML)Machine learning is the practice of writing programs that learn a function from data, via a loss function and an optimizer, instead of following hand-written rules. turns "write a program that does X" into "write a program that learns to do X from examples." Concretely, three ingredients:
- A model: a function
f(x; θ)that maps an inputxto a prediction, controlled by a set of adjustable numbersθ(the parameters, or weights). Everything from a straight line (y = mx + b, whereθ = {m, b}) to a 70-billion-parameter transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. fits this shape. - A loss functionLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training.: a single number measuring how wrong the model's predictions are on some data. Lower is better, by construction.
- An optimizer: an algorithm that nudges
θto make the loss smaller. In deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones., this is almost always gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. (below).
Training a model just means: repeatedly measure the loss, compute how to
adjust θ to reduce it, and adjust. That's it — that loop, run millions of
times, is the entire "learning" in machine learning.
Supervised vs. unsupervised learning
The main axis along which ML problems differ is what data you have:
- Supervised learning: your data comes as
(input, correct output)pairs — a photo labeled "cat," an email labeled "spam," an English sentence paired with its French translation. The loss function directly compares the model's prediction to the known correct answer. Most of the practical systems in this course (image classifiers, translation, next-token prediction) are supervised, or a variant of it. - Unsupervised learning: your data has no labels — just raw examples — and the goal is to find structure in it anyway: group similar items together (clustering), compress data while preserving its structure (dimensionality reduction), or model the distribution the data came from.
A useful middle case you'll see constantly in modern AI is self-supervised learning: you manufacture labels from the data itself. Language model pretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability. is the canonical example — take a sentence, hide the next word, and use that hidden word as the "label." No human annotated anything; the supervision signal comes from the structure already present in raw text. This is why LLMs can pretrain on the entire internet without anyone hand-labeling it.
Loss functions: turning "wrong" into a number
A loss function has to translate "how bad was this prediction" into a differentiable number (see below for why differentiable matters). Two you will see everywhere:
- Mean squared error (MSE), for predicting continuous numbers (house prices, temperatures): average the squared difference between prediction and truth. Squaring punishes large errors disproportionately and keeps the loss positive.
- Cross-entropy loss, for classification (is this a cat, a dog, or neither?): penalizes a model heavily for being confidently wrong, and only lightly for being unsure. This is the loss behind next-token prediction in language models — at every position, "classify" which token comes next out of the entire vocabulary.
The choice of loss function encodes what you actually want the model to optimize for — a frequent source of real-world ML bugs is a loss function that's easy to write down but doesn't quite match the behavior you want.
For examples with true values and predictions , mean squared error is:
For classification over classes, where is 1 for the correct class and 0 otherwise, and is the model's predicted probability:
Because only the correct class survives the inner sum, this reduces to . That logarithm is what makes confident mistakes so expensive: predicting 0.01 for the correct class costs , while predicting 0.5 costs only . As the predicted probability approaches zero the loss approaches infinity, so the optimizer will do almost anything to avoid being confidently wrong.
Gradient descent
Once you have a loss, you need a way to change θ to reduce it. Gradient
descent's insight: the gradient of the loss with respect to θ (its
multivariable derivative) points in the direction of steepest increase.
So to decrease the loss, step in the opposite direction:
θ ← θ - η · ∇L(θ)Where ∇L(θ) is the gradient of the loss L at the current parameters,
and η (eta) is the learning rate — how big a step to take. This
single update rule, applied repeatedly, is what "training" means for
essentially every model in this course, from linear regression to GPT-4
scale transformers.
Click anywhere on the surface to drop a starting point, then run gradient descent on f(x, y) = x² + 2y². Watch how the learning rate changes the path.
Two practical wrinkles worth knowing by name, since you'll hit both:
- Learning rate matters a lot. Too small and training crawls; too large and updates overshoot the minimum and the loss diverges instead of decreasing.
- Stochastic gradient descent (SGD): computing the exact gradient requires averaging over your entire dataset, which is too slow at scale. In practice, you estimate the gradient from a small random batch of examples at a time (a mini-batch) and update immediately — noisier per step, but far more updates per unit of compute, and it works better in practice than it has any right to.
Modern training almost never uses plain SGD directly — optimizers like Adam adapt the effective learning rate per-parameter based on the history of past gradients — but they're all still gradient descent at heart: compute a gradient, step against it, repeat. Optimization & Training Dynamics covers Adam's actual mechanics, why warmup exists, and what batch size trades off at scale, in depth.
Suppose 1 million training examples and one gradient step per full pass. At a generous 10,000 examples/second, one step takes 100 seconds. Training needs tens of thousands of steps, so a full-batch run would take years.
Use mini-batches of 32 instead and the same 100 seconds buys you ~31,000 steps. Each step is noisier, but you get four orders of magnitude more of them — and averaging over 32 samples already estimates the gradient direction well enough to make progress.
Gradient descent minimises by repeatedly applying:
where is the vector of partial derivatives — one per parameter.
The gradient points along the direction of steepest increase, which is why the update subtracts it. The learning rate controls step size, and the failure modes sit on either side of it: for a locally quadratic loss with curvature , updates diverge once , while very small converges at a rate proportional to — correct, but arbitrarily slow.
Stochastic gradient descent replaces the full-dataset gradient with an estimate from a mini-batch :
This estimate is unbiased — correct on average — and its variance falls as . So quadrupling the batch size only halves the gradient noise, which is exactly why small batches win on a compute budget.
Generalization: the goal was never to minimize training loss
Everything above describes how to drive the loss down on the data you're training on. That is not actually the goal. The goal is a model that performs well on new data it has never seen — a spam filter has to work on tomorrow's emails, not just the ones in its training set. A model that only memorizes its training data is worthless in production even if its training loss is zero.
This gap — performance on training data vs. performance on new data — is the central practical concern of applied ML, and it's why you never just train on all your data and call it done.
Train / validation / test splits
Before training, data is split into (at least) three disjoint sets:
- Training set: what the model actually learns from — gradient descent
updates
θusing only this data. - Validation set: held out from training, used to check performance during development — comparing architectures, tuning the learning rate, deciding when to stop training. Because you look at validation performance repeatedly while making these decisions, you can unintentionally overfit to the validation set too, just more slowly.
- Test set: touched exactly once, at the very end, to report final performance. If you tune anything based on test-set results, it stops being a trustworthy estimate of real-world performance — it's now something you optimized against, not an independent check.
Overfitting and underfitting
- OverfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data.: the model fits the training data very well — including its noise and idiosyncrasies — but that doesn't transfer to new data. Telltale sign: training loss keeps dropping while validation loss stops improving or gets worse. A model with enough parameters can, in principle, memorize its entire training set without learning any generalizable pattern at all.
- Underfitting: the model is too simple (or undertrained) to capture even the patterns present in the training data — both training and validation loss stay high.
Reading the two curves together is enough to diagnose which one you have:
The bias-variance tradeoffBias-Variance TradeoffThe bias-variance tradeoff describes the balance between a model too simple to fit the data (high bias) and one that overfits its training data's noise (high variance). is the formal name for this balance: a model that's too simple has high bias (systematically wrong, underfitting); a model that's too flexible has high variance (fits the specific quirks of whatever training data it happened to see, overfitting). The practical goal is the sweet spot between the two, not the maximum-capacity model.
For squared-error loss, the expected error on an unseen point decomposes exactly into three terms:
The expectations run over different training sets drawn from the same distribution. Bias is how far the average model is from the truth; variance is how much the model moves as the training data changes; is label noise, and no model can beat it.
Two consequences worth internalising. First, since only two of three terms are reducible, a non-zero validation loss is not automatically a bug. Second, the classical U-shaped curve this predicts is not the whole story: very large neural networks enter a double descent regime where test error falls again past the interpolation point, which is a large part of why scaling up keeps working.
Regularization: fighting overfitting on purpose
RegularizationRegularizationRegularization is any technique that trades some training-data fit for better generalization, fighting overfitting on purpose. is any technique that trades a bit of training-data fit for better generalization. Common ones you'll see by name throughout this course:
- Weight decay (L2 regularization): adds a penalty to the loss for large weight values, discouraging the model from relying too heavily on any single feature.
- Dropout: during training, randomly zero out a fraction of neurons' outputs on each forward pass, forcing the network to not depend too heavily on any specific neuron — effectively training a large ensemble of overlapping sub-networks that share weights.
- Early stopping: simply stop training when validation loss stops improving, even if training loss would keep dropping — the simplest regularizer of all.
- More data: not a technique so much as the underlying fix — the less data relative to model capacity, the more a model can overfit; more (varied) data directly shrinks the overfitting gap.
Evaluation metrics: loss isn't the whole story
The loss function is what gradient descent optimizes, but it's often not the number a human actually cares about. For a spam classifier, a few metrics computed on the validation/test set are usually more meaningful than the raw cross-entropy value:
- Accuracy: fraction of predictions that were correct — intuitive, but misleading on imbalanced data (a filter that always predicts "not spam" gets 99% accuracy if only 1% of email is spam).
- Precision: of everything the model flagged as positive, what fraction actually was? High precision means few false alarms.
- Recall: of everything that actually was positive, what fraction did the model catch? High recall means few misses.
- F1 score: the harmonic mean of precision and recall, used when you want one number balancing both.
| Metric | Question it answers | Fails when |
|---|---|---|
| Accuracy | How often is the model right? | Classes are imbalanced |
| Precision | When it says yes, is it right? | You care about missed cases |
| Recall | Does it catch everything? | False alarms are expensive |
| F1 | Balance of precision and recall | The two have different costs |
1% of email is spam. A model that predicts "not spam" for every single message scores:
- Accuracy: 99% — looks excellent
- Recall: 0% — it catches no spam at all
- Precision: undefined — it never flags anything
Same model, three wildly different verdicts. This is why you never report accuracy alone on imbalanced data.
Precision and recall trade off against each other (a model can trivially get 100% recall by flagging everything, tanking precision) — which one matters more depends entirely on the cost of a false positive vs. a false negative for the specific application, a judgment call the loss function alone can't make for you.
Why this matters going forward
The next lesson steps back to the probability and statistics underneath
loss functions themselves — why cross-entropy and MSE are the "correct"
choices, not arbitrary ones. After that,
Neural Networks & Backprop covers
how you efficiently compute ∇L(θ) when θ has millions or billions
of parameters arranged in layers — that algorithm is called
backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network., and it's the single most load-bearing idea in deep
learning. Everything after that — CNNsCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently., transformers, LLMs — is a choice of
what function f(x; θ) looks like. The loss-and-gradient-descent loop
underneath stays exactly the same.
History & Landscape
Symbolic AI to expert systems to statistical ML to deep learning to the LLM era
Optimization & Training Dynamics
Why plain gradient descent isn't what actually trains modern models — Adam's mechanics, why warmup exists, how learning-rate schedules are shaped, and what batch size actually trades off at scale