Policy
A policy is the strategy a reinforcement-learning agent learns for choosing actions — the thing being trained, analogous to a model in supervised learning.
A policy maps states to actions — it's the thing a
reinforcement learning agent is trying
to learn, analogous to the model f(x; θ) in supervised learning, except
here the "prediction" is an action to take rather than a label.
How it works
Two broad families learn a policy differently. Value-based methods (like Q-learningQ-LearningQ-learning learns the expected future reward of taking a specific action in a specific state, updating its estimate incrementally from every action the agent actually takes.) learn a value function first — how much future reward to expect from a state or action — and derive a policy indirectly, by always choosing the action with the highest estimated value. Policy gradient methods skip that intermediate step and directly adjust the policy's parameters to make higher-reward actions more likely, using gradient ascent on an objective built from the rewards actually received.
Either way, a policy has to balance exploitation (taking the action it currently believes is best) against exploration (trying something uncertain to potentially discover something better) — a tension with no equivalent in supervised learning, where the training data is fixed rather than shaped by the model's own past decisions.
When it breaks
- A stochastic policy is a probability distribution over actions, not a single choice. This is deliberate — sampling from it is what makes exploration possible at all, and policy-gradient methods specifically need the ability to compute a probability for the action taken.
- Large policy updates can catastrophically break a working policy. Because the policy also determines what data gets collected next, an update that moves too far can send the agent into a state distribution it never learns to recover from — the specific problem PPOPPO (Proximal Policy Optimization)PPO is a policy-gradient reinforcement-learning algorithm that takes conservative, clipped update steps, and is the algorithm most commonly used to optimize LLMs during RLHF. was designed to prevent.
- A policy trained on a proxy reward optimizes the proxy, not the intent behind it. See rewardRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one. hacking.
See also: MDPMDP (Markov Decision Process)An MDP formalizes the reinforcement-learning setup — an agent observing a state, choosing an action, and receiving a reward, with no dataset of correct actions to imitate., Q-LearningQ-LearningQ-learning learns the expected future reward of taking a specific action in a specific state, updating its estimate incrementally from every action the agent actually takes., PPOPPO (Proximal Policy Optimization)PPO is a policy-gradient reinforcement-learning algorithm that takes conservative, clipped update steps, and is the algorithm most commonly used to optimize LLMs during RLHF.
Learn more: Reinforcement Learning
Mentioned in
Lessons where this comes up in context.
- AI Safety & AlignmentWhether a model's own objectives match what we actually want, independent of any attacker — specification gaming, outer vs. inner alignment, why RLHF isn't a complete answer, and scalable oversight
- Reinforcement LearningMDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF
Reward
A reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one.
Q-Learning
Q-learning learns the expected future reward of taking a specific action in a specific state, updating its estimate incrementally from every action the agent actually takes.