Reward
A reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one.
A reward is the number a reinforcement learning agent receives from its environment after taking an action, indicating how good or bad that action was. It's the entire learning signal in RL — there's no per-example correct answer to imitate, only this scalar, which is frequently delayed until long after the action that earned it.
How it works
An agent's objective is to maximize total reward over time (the return), not the immediate reward from a single action — a discount factor typically weights near-term reward more heavily than distant reward, both to keep the total finite over long or infinite horizons and to set an effective planning horizon. Because a single delayed reward has to explain an entire sequence of actions, RL algorithms generally learn a value function — an estimate of how much future reward to expect from a state — as an intermediate signal that turns one sparse, delayed number into something closer to per-step feedback.
When it breaks
- A high-scoring agent has learned to make the reward go up, not necessarily to do the intended task. This is reward hacking: agents reliably find loopholes a reward's designer didn't anticipate — collecting points by circling a pickup instead of finishing a race, or exploiting a physics bug rather than actually walking.
- The reward is only as good as its design. Specifying a reward that actually captures the intended goal, without leaving exploitable gaps, is a genuinely hard design problem — a loss function that's easy to write down often isn't the behavior anyone wanted.
- Sparse reward makes credit assignment expensive. A single win/loss signal at the end of a long game carries very little information per decision, which is why RL agents are typically trained on millions to billions of environment steps.
See also: MDPMDP (Markov Decision Process)An MDP formalizes the reinforcement-learning setup — an agent observing a state, choosing an action, and receiving a reward, with no dataset of correct actions to imitate., PolicyPolicyA policy is the strategy a reinforcement-learning agent learns for choosing actions — the thing being trained, analogous to a model in supervised learning., PPOPPO (Proximal Policy Optimization)PPO is a policy-gradient reinforcement-learning algorithm that takes conservative, clipped update steps, and is the algorithm most commonly used to optimize LLMs during RLHF.
Learn more: Reinforcement Learning
Mentioned in
Lessons where this comes up in context.
- AI Safety & AlignmentWhether a model's own objectives match what we actually want, independent of any attacker — specification gaming, outer vs. inner alignment, why RLHF isn't a complete answer, and scalable oversight
- Evaluation & BenchmarksHow models are actually scored — benchmark suites, leaderboards, eval methodology, and why a high benchmark score doesn't guarantee good real-world performance
- Generative ModelsGANs, VAEs, and diffusion models — how AI generates new images, audio, and video, as opposed to classifying or understanding existing content
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- Multimodal ModelsHow a single model handles more than one kind of input — CLIP's shared embedding space, how vision gets fed into a language model as tokens, and what actually breaks
- Probability & Statistics FoundationsDistributions, Bayes' theorem, and maximum likelihood estimation — the math that loss functions and uncertainty in ML are actually built on
- Reinforcement LearningMDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF
MDP (Markov Decision Process)
An MDP formalizes the reinforcement-learning setup — an agent observing a state, choosing an action, and receiving a reward, with no dataset of correct actions to imitate.
Policy
A policy is the strategy a reinforcement-learning agent learns for choosing actions — the thing being trained, analogous to a model in supervised learning.