RLHF (Reinforcement Learning from Human Feedback)
RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward.
RLHF (Reinforcement Learning from Human Feedback) aligns an LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. with human preferences beyond what fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions. alone can teach. Steps: collect human rankings of multiple model responses to the same prompt, train a reward model to predict those rankings, then use reinforcement learning to optimize the LLM to produce responses that score highly under that reward model. DPO (Direct Preference Optimization) is a newer, simpler alternative that skips the separate reward model and RL loop.
How it works
Why train a separate reward model instead of optimizing on preferences directly?
Three stages, each producing an artifact:
- Preference collection. Sample several responses to the same prompt and have annotators rank them. The data is pairwise, not scored, because relative judgments are far more consistent between raters.
- Reward model. Take a copy of the
pretrainedPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability. model, attach a scalar head, and
train it with a ranking lossLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. so chosen
responses score above rejected ones. It maps a
(prompt, response)pair to one number. - Policy optimization. Run PPO: sample responses from the current policy, score them with the reward model, and take gradient steps that raise expected reward. A KL penalty against the starting model keeps the policy from drifting into text that scores well and reads badly.
DPODPO (Direct Preference Optimization)DPO aligns an LLM to human preferences directly from ranked response pairs, without the separate reward model and reinforcement learning loop RLHF requires. collapses the last two stages into a single loss.
When it breaks
- Reward hacking. The policy finds regions where the reward model is wrong — hedging, padding, flattery — and exploits them. Measured reward rises while actual quality falls.
- The KL penalty is a tightrope. Too weak and the policy degenerates; too strong and preferences never take hold.
- Rater habits become model behavior. Annotators favor long, confident answers, so the model learns verbosity and overconfidence alongside helpfulness.
- Heavy machinery. PPO keeps policy, reference, reward model, and value model in play at once, making runs expensive and far less stable than supervised training.
See also: Fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions., HallucinationHallucinationA hallucination is a confidently stated, fluent LLM output that is factually wrong, a direct consequence of models being trained to produce plausible text rather than verified facts.
Learn more: LLMs · Paper: InstructGPT
Mentioned in
Lessons where this comes up in context.
- AI Safety & AlignmentWhether a model's own objectives match what we actually want, independent of any attacker — specification gaming, outer vs. inner alignment, why RLHF isn't a complete answer, and scalable oversight
- AI SecurityAdversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
- Applied & Agentic SystemsHow prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
- Evaluation & BenchmarksHow models are actually scored — benchmark suites, leaderboards, eval methodology, and why a high benchmark score doesn't guarantee good real-world performance
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- Reinforcement LearningMDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF
LoRA (Low-Rank Adaptation)
LoRA is a parameter-efficient fine-tuning method that trains a small number of additional low-rank parameters instead of updating an entire model's weights.
DPO (Direct Preference Optimization)
DPO aligns an LLM to human preferences directly from ranked response pairs, without the separate reward model and reinforcement learning loop RLHF requires.