PPO (Proximal Policy Optimization)
PPO is a policy-gradient reinforcement-learning algorithm that takes conservative, clipped update steps, and is the algorithm most commonly used to optimize LLMs during RLHF.
PPO (Proximal Policy Optimization) is a widely used policy gradient algorithm: instead of learning a value function first, it directly adjusts a policyPolicyA policy is the strategy a reinforcement-learning agent learns for choosing actions — the thing being trained, analogous to a model in supervised learning.'s parameters to make higher-reward actions more likely, using gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. on an objective built from the rewards received.
How it works
Policy gradient methods have a specific failure mode PPO is designed around: a large update can catastrophically break a policy that was working, because the policy also determines what data gets collected next — an update that moves too far can send the agent into a part of the state space it never recovers from. PPO fixes this with a clipped objective that refuses to let the probability ratio between the new and old policy stray far from 1, trading some update speed for training stability.
PPO is specifically the RL algorithm most commonly used inside RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward.: an LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. is the policy, generating a response is taking an action, and a learned reward model's score is the reward. Everything about PPO — stable, conservative, incremental updates — is why it, rather than a less constrained policy-gradient method, is the standard choice for that step.
When it breaks
- Clipping bounds each individual update, not the final result. A separate KL-divergence penalty against a frozen reference model is typically added on top, specifically in RLHF, to bound how far the cumulative policy drift can go — clipping alone doesn't prevent slow drift across many small steps.
- Still vulnerable to reward hacking. PPO optimizes exactly the reward signal it's given; if that reward (e.g. a learned reward model) is an imperfect proxy for what's actually wanted, PPO will find and exploit the gap just as readily as any other RL method.
- Sample-inefficient compared to value-based methods on tasks where a discrete action space and a tabular or DQN-style approach fit naturally — PPO's generality costs more environment interactions to reach the same performance in those cases.
See also: PolicyPolicyA policy is the strategy a reinforcement-learning agent learns for choosing actions — the thing being trained, analogous to a model in supervised learning., RewardRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one., RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward.
Learn more: Reinforcement Learning
Mentioned in
Lessons where this comes up in context.
Q-Learning
Q-learning learns the expected future reward of taking a specific action in a specific state, updating its estimate incrementally from every action the agent actually takes.
Inference
Inference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering.