Foundations

ML Fundamentals

Supervised/unsupervised learning, loss functions, gradient descent

Every idea in the rest of this course — neural networks, attention, LLMs — is a variation on one core recipe. Learn this recipe well and the rest is mostly "same idea, fancier function."

The recipe: model, loss, optimize

Machine learning turns "write a program that does X" into "write a program that learns to do X from examples." Concretely, three ingredients:

  1. A model: a function f(x; θ) that maps an input x to a prediction, controlled by a set of adjustable numbers θ (the parameters, or weights). Everything from a straight line (y = mx + b, where θ = {m, b}) to a 70-billion-parameter transformer fits this shape.
  2. A loss function: a single number measuring how wrong the model's predictions are on some data. Lower is better, by construction.
  3. An optimizer: an algorithm that nudges θ to make the loss smaller. In deep learning, this is almost always gradient descent (below).

Training a model just means: repeatedly measure the loss, compute how to adjust θ to reduce it, and adjust. That's it — that loop, run millions of times, is the entire "learning" in machine learning.

Supervised vs. unsupervised learning

The main axis along which ML problems differ is what data you have:

  • Supervised learning: your data comes as (input, correct output) pairs — a photo labeled "cat," an email labeled "spam," an English sentence paired with its French translation. The loss function directly compares the model's prediction to the known correct answer. Most of the practical systems in this course (image classifiers, translation, next-token prediction) are supervised, or a variant of it.
  • Unsupervised learning: your data has no labels — just raw examples — and the goal is to find structure in it anyway: group similar items together (clustering), compress data while preserving its structure (dimensionality reduction), or model the distribution the data came from.

A useful middle case you'll see constantly in modern AI is self-supervised learning: you manufacture labels from the data itself. Language model pretraining is the canonical example — take a sentence, hide the next word, and use that hidden word as the "label." No human annotated anything; the supervision signal comes from the structure already present in raw text. This is why LLMs can pretrain on the entire internet without anyone hand-labeling it.

Loss functions: turning "wrong" into a number

A loss function has to translate "how bad was this prediction" into a differentiable number (see below for why differentiable matters). Two you will see everywhere:

  • Mean squared error (MSE), for predicting continuous numbers (house prices, temperatures): average the squared difference between prediction and truth. Squaring punishes large errors disproportionately and keeps the loss positive.
  • Cross-entropy loss, for classification (is this a cat, a dog, or neither?): penalizes a model heavily for being confidently wrong, and only lightly for being unsure. This is the loss behind next-token prediction in language models — at every position, "classify" which token comes next out of the entire vocabulary.

The choice of loss function encodes what you actually want the model to optimize for — a frequent source of real-world ML bugs is a loss function that's easy to write down but doesn't quite match the behavior you want.

Gradient descent

Once you have a loss, you need a way to change θ to reduce it. Gradient descent's insight: the gradient of the loss with respect to θ (its multivariable derivative) points in the direction of steepest increase. So to decrease the loss, step in the opposite direction:

θ ← θ - η · ∇L(θ)

Where ∇L(θ) is the gradient of the loss L at the current parameters, and η (eta) is the learning rate — how big a step to take. This single update rule, applied repeatedly, is what "training" means for essentially every model in this course, from linear regression to GPT-4 scale transformers.

Click anywhere on the surface to drop a starting point, then run gradient descent on f(x, y) = x² + 2y². Watch how the learning rate changes the path.

step: 0loss: 10.8800x: 2.400y: 1.600

Two practical wrinkles worth knowing by name, since you'll hit both:

  • Learning rate matters a lot. Too small and training crawls; too large and updates overshoot the minimum and the loss diverges instead of decreasing.
  • Stochastic gradient descent (SGD): computing the exact gradient requires averaging over your entire dataset, which is too slow at scale. In practice, you estimate the gradient from a small random batch of examples at a time (a mini-batch) and update immediately — noisier per step, but far more updates per unit of compute, and it works better in practice than it has any right to.

Modern training almost never uses plain SGD directly — optimizers like Adam adapt the effective learning rate per-parameter based on the history of past gradients — but they're all still gradient descent at heart: compute a gradient, step against it, repeat. Optimization & Training Dynamics covers Adam's actual mechanics, why warmup exists, and what batch size trades off at scale, in depth.

Why mini-batches exist

Suppose 1 million training examples and one gradient step per full pass. At a generous 10,000 examples/second, one step takes 100 seconds. Training needs tens of thousands of steps, so a full-batch run would take years.

Use mini-batches of 32 instead and the same 100 seconds buys you ~31,000 steps. Each step is noisier, but you get four orders of magnitude more of them — and averaging over 32 samples already estimates the gradient direction well enough to make progress.

Generalization: the goal was never to minimize training loss

Everything above describes how to drive the loss down on the data you're training on. That is not actually the goal. The goal is a model that performs well on new data it has never seen — a spam filter has to work on tomorrow's emails, not just the ones in its training set. A model that only memorizes its training data is worthless in production even if its training loss is zero.

This gap — performance on training data vs. performance on new data — is the central practical concern of applied ML, and it's why you never just train on all your data and call it done.

Train / validation / test splits

Before training, data is split into (at least) three disjoint sets:

  • Training set: what the model actually learns from — gradient descent updates θ using only this data.
  • Validation set: held out from training, used to check performance during development — comparing architectures, tuning the learning rate, deciding when to stop training. Because you look at validation performance repeatedly while making these decisions, you can unintentionally overfit to the validation set too, just more slowly.
  • Test set: touched exactly once, at the very end, to report final performance. If you tune anything based on test-set results, it stops being a trustworthy estimate of real-world performance — it's now something you optimized against, not an independent check.

Overfitting and underfitting

  • Overfitting: the model fits the training data very well — including its noise and idiosyncrasies — but that doesn't transfer to new data. Telltale sign: training loss keeps dropping while validation loss stops improving or gets worse. A model with enough parameters can, in principle, memorize its entire training set without learning any generalizable pattern at all.
  • Underfitting: the model is too simple (or undertrained) to capture even the patterns present in the training data — both training and validation loss stay high.

Reading the two curves together is enough to diagnose which one you have:

The bias-variance tradeoff is the formal name for this balance: a model that's too simple has high bias (systematically wrong, underfitting); a model that's too flexible has high variance (fits the specific quirks of whatever training data it happened to see, overfitting). The practical goal is the sweet spot between the two, not the maximum-capacity model.

Regularization: fighting overfitting on purpose

Regularization is any technique that trades a bit of training-data fit for better generalization. Common ones you'll see by name throughout this course:

  • Weight decay (L2 regularization): adds a penalty to the loss for large weight values, discouraging the model from relying too heavily on any single feature.
  • Dropout: during training, randomly zero out a fraction of neurons' outputs on each forward pass, forcing the network to not depend too heavily on any specific neuron — effectively training a large ensemble of overlapping sub-networks that share weights.
  • Early stopping: simply stop training when validation loss stops improving, even if training loss would keep dropping — the simplest regularizer of all.
  • More data: not a technique so much as the underlying fix — the less data relative to model capacity, the more a model can overfit; more (varied) data directly shrinks the overfitting gap.

Evaluation metrics: loss isn't the whole story

The loss function is what gradient descent optimizes, but it's often not the number a human actually cares about. For a spam classifier, a few metrics computed on the validation/test set are usually more meaningful than the raw cross-entropy value:

  • Accuracy: fraction of predictions that were correct — intuitive, but misleading on imbalanced data (a filter that always predicts "not spam" gets 99% accuracy if only 1% of email is spam).
  • Precision: of everything the model flagged as positive, what fraction actually was? High precision means few false alarms.
  • Recall: of everything that actually was positive, what fraction did the model catch? High recall means few misses.
  • F1 score: the harmonic mean of precision and recall, used when you want one number balancing both.
MetricQuestion it answersFails when
AccuracyHow often is the model right?Classes are imbalanced
PrecisionWhen it says yes, is it right?You care about missed cases
RecallDoes it catch everything?False alarms are expensive
F1Balance of precision and recallThe two have different costs
The imbalanced-data trap

1% of email is spam. A model that predicts "not spam" for every single message scores:

  • Accuracy: 99% — looks excellent
  • Recall: 0% — it catches no spam at all
  • Precision: undefined — it never flags anything

Same model, three wildly different verdicts. This is why you never report accuracy alone on imbalanced data.

Precision and recall trade off against each other (a model can trivially get 100% recall by flagging everything, tanking precision) — which one matters more depends entirely on the cost of a false positive vs. a false negative for the specific application, a judgment call the loss function alone can't make for you.

Why this matters going forward

The next lesson steps back to the probability and statistics underneath loss functions themselves — why cross-entropy and MSE are the "correct" choices, not arbitrary ones. After that, Neural Networks & Backprop covers how you efficiently compute ∇L(θ) when θ has millions or billions of parameters arranged in layers — that algorithm is called backpropagation, and it's the single most load-bearing idea in deep learning. Everything after that — CNNs, transformers, LLMs — is a choice of what function f(x; θ) looks like. The loss-and-gradient-descent loop underneath stays exactly the same.

On this page