Reinforcement Learning
MDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF
ML Fundamentals introduced supervised and unsupervised learning as the two main paradigms. Reinforcement learning (RL) is the third: instead of learning from labeled examples or finding structure in unlabeled data, an RL agentAgentAn agent puts an LLM in a loop with tool access, letting it decide autonomously which tools to call and in what order to accomplish a multi-step goal. learns by acting in an environment and receiving feedback about how good its actions were. RLHF leaned on this paradigm without explaining it — this lesson fills that gap properly.
The setup: agent, environment, reward
RL problems share a common structure, formalized as a Markov Decision Process (MDPMDP (Markov Decision Process)An MDP formalizes the reinforcement-learning setup — an agent observing a state, choosing an action, and receiving a reward, with no dataset of correct actions to imitate.):
- An agent observes a state of the environment and chooses an action.
- The environment transitions to a new state and returns a rewardRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one. — a number indicating how good or bad that action was.
- The agent's goal is to learn a policyPolicyA policy is the strategy a reinforcement-learning agent learns for choosing actions — the thing being trained, analogous to a model in supervised learning. — a strategy for choosing actions — that maximizes its total reward over time, not just the immediate reward from one action.
What's the difference between reinforcement learning and supervised learning?
That last point is what makes RL genuinely different from supervised learning: there's no dataset of "correct" actions to imitate. An action that looks bad immediately (sacrificing a chess piece) can be the right choice if it leads to much better rewards later (winning the game). The agent has to learn this from trial and error, via the reward signal alone.
The whole field is built on one loop, repeated until the policy stops improving:
Compared side by side with supervised learning, almost every practical difficulty in RL traces back to the shape of that loop:
| Supervised learning | Reinforcement learning | |
|---|---|---|
| Signal | Correct answer per example | Scalar reward, often delayed |
| Data source | Fixed dataset | Generated by the agent itself |
| Stability | Loss usually descends smoothly | Policy changes change the data |
The middle row is the one that bites. In supervised learning the dataset sits still while you train; in RL, improving the policy changes which states the agent visits, which changes the data it learns from next — a feedback loop with no equivalent in the previous lessons.
Policies and value functions
A policy (often written π) maps states to actions — it's the thing
being learned, analogous to the model f(x; θ) from
ML Fundamentals, except here the "prediction" is
an action to take, not a label.
Rather than learning a policy directly, many RL algorithms instead learn a value function: an estimate of how much total future reward to expect from a given state (or from taking a specific action in a state), if the agent behaves well from then on. The intuition: if you know how valuable every state is, choosing a good action is easy — just move toward higher-value states. This is a genuinely different structure from the loss functionsLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. covered so far: rather than measuring "how wrong is this single prediction," a value function estimates "how good is this situation, accounting for everything that could happen afterward."
"Total future reward" needs to be made precise. The return from time is the discounted sum of every reward that follows:
with discount factor . The value function of a policy is the expected return from a state, and the action-value (or Q) function conditions additionally on the first action:
Both satisfy a recursive Bellman equation, which is what makes them learnable one step at a time rather than by simulating to the end of time:
The Q-learningQ-LearningQ-learning learns the expected future reward of taking a specific action in a specific state, updating its estimate incrementally from every action the agent actually takes. update in the next section is exactly this equation turned into an incremental estimator, with in place of so it converges toward the optimal policy rather than the current one.
The discount factor buys two things. Mathematically, it keeps the sum finite: if every reward is bounded by , then , so an infinite-horizon task still has a well-defined value. Practically, it sets an effective planning horizon of roughly steps — means the agent is reasoning about the next hundred or so decisions, and rewards much further out barely register.
A chess game runs ~40 moves per side and ends with a single number: +1, 0, or −1. Compare the bookkeeping:
- Supervised: 40 decisions would come with 40 labels — one per move.
- RL: 40 decisions share ~1.6 bits of feedback, total.
So the signal per decision is roughly 1/40th of a label, and it is the same number for the brilliant move and the blunder in the same game. The only way out is volume: play the game enough times that the moves systematically preceding wins can be told apart from the moves that happened to be nearby. That's why RL agents are typically trained on millions to billions of environment steps, while a supervised image classifier learns from a dataset of perhaps a million labeled examples.
Q-learning: learning values from experience
Q-learning is a foundational RL algorithm that learns a Q-function — the expected total future reward of taking a specific action in a specific state, then acting optimally afterward. It updates its estimate after every action taken, using the actual reward received plus its current estimate of the next state's value:
Q(state, action) ← Q(state, action) + α · [reward + γ · max(Q(next_state, ·)) − Q(state, action)]Where α is a learning rate (same role as in
gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient.) and γ (gamma) is a
discount factor — a number slightly less than 1 that makes future
rewards count slightly less than immediate ones, so the agent doesn't
value a reward infinitely far in the future the same as one right now.
Deep Q-Networks (DQN) replace the lookup-table version of Q with a
neural networkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. — the same substitution
covered throughout this course, swapping a simple representation for a
learned one that can generalize across states it hasn't seen exactly
before. This was the approach behind early landmark results like
learning to play Atari games directly from raw pixels.
The lookup-table version is small enough to train live, right here. Click "Train ×50" a few times and watch a policy emerge — arrows pointing toward the flag, curving around the trap:
Untrained — every cell starts at Q = 0, so the arrows would be meaningless. Train a few episodes to see a policy emerge.
Policy gradients: learning the policy directly
An alternative family of methods, policy gradient methods, skip value functions and directly adjust a policy's parameters to make higher-reward actions more likely — using gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. on an objective built from the rewards received, in the same spirit as (though mechanically different from) the loss-minimization loop from ML Fundamentals. PPOPPO (Proximal Policy Optimization)PPO is a policy-gradient reinforcement-learning algorithm that takes conservative, clipped update steps, and is the algorithm most commonly used to optimize LLMs during RLHF. (Proximal Policy Optimization) is a widely used policy gradient method, specifically designed to take stable, conservative update steps — large policy updates in RL can catastrophically break a policy that was working, so limiting how much any single update is allowed to change it improves training stability.
Write the policy as — a probability distributionProbability DistributionA probability distribution assigns a likelihood to each possible value of a random variable — the object every loss function is secretly built to measure the fit of. over actions, parameterised by . The objective is expected return , and the policy gradient theorem gives its gradient a form you can estimate from sampled trajectories:
Read it as an instruction: increase the log-probability of each action, weighted by how good the outcome was. Gradient ascent on this pushes probability mass toward actions that preceded high return, without ever needing to know what the correct action would have been.
The naive estimator is extremely noisy, because mixes the quality of the action with the quality of the situation the agent happened to be in. Subtracting a baseline that depends only on the state fixes this, and does so for free: since for any fixed , the subtraction leaves the gradient unbiased while shrinking its variance. Taking gives the advantage:
Now the weight is how much better than average the action was, so an action is only reinforced if it beat the expectation for that state — not merely because the state was a good one to be in. PPO builds on exactly this, adding a clipped objective that refuses to let the probability ratio between the new and old policy stray far from 1.
This is the algorithm behind RLHF
LLMs described RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward. as "using reinforcement learning to optimize the LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. against a learned reward model." PPO is specifically the RL algorithm most commonly used for that optimization step: the LLM itself is the policy, generating a response is taking an action, and the reward model's score is the reward. Everything in this lesson — a policy, a reward signal, stable incremental updates — maps directly onto that pipeline.
The mapping is almost one-to-one, with a learned stand-in for the environment:
- State: the prompt plus the tokens generated so far.
- Action: the next token (or, viewed coarsely, the whole response).
- Policy: the language model itself, .
- Environment: a reward model trained on human preference comparisons, since there is no simulator that can score an essay.
- Reward: for prompt and response , delivered once at the end of the sequence.
Optimising that reward alone would be a disaster: the policy would drift into whatever text maximises the reward model's score, including text the reward model was never trained on and scores wrongly. RLHF therefore penalises divergence from the frozen pre-RL reference model :
The coefficient is the dial between the two failure modes: too small and the policy exploits the reward model and degenerates; too large and the model barely changes from where it started. Note that the KL term plays a different role from PPO's clipping — clipping stabilises each individual update, while the KL penalty bounds how far the final policy may end up from the reference.
Exploration vs. exploitation
A tension unique to RL: should the agent take the action it currently believes is best (exploitation), or try something uncertain to potentially discover something better (exploration)? An agent that only exploits can get stuck with a mediocre policy it never learns to improve on; an agent that only explores never capitalizes on what it's learned. Most RL algorithms build in some explicit mechanism to balance the two — this problem has no analog in supervised learning, where the training data is fixed rather than shaped by the model's own past decisions.
Recap and what's next
Reinforcement learning learns from a reward signal generated by acting in an environment, rather than from labeled examples — a genuinely different problem structure from supervised learning, requiring policies, value functions, and a balance between exploration and exploitation. Its practical relevance to this course is direct: it's the actual algorithm family underneath RLHF's alignmentAlignmentAlignment is whether a trained model's actual objective and behavior match what its designers intended, split into outer and inner alignment. step. The next lesson returns to inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. — what happens when you take a trained model, RL-tuned or otherwise, and actually run it.