DPO (Direct Preference Optimization)
DPO aligns an LLM to human preferences directly from ranked response pairs, without the separate reward model and reinforcement learning loop RLHF requires.
DPO (Direct Preference Optimization) is a simpler alternative to RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward. for aligning an LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. with human preferences. Instead of training a separate reward model and then running reinforcement learning against it, DPO optimizes the LLM directly on preference-ranked response pairs, using a loss function derived to reach the same optimum RLHF would. It's become popular for being simpler to implement and tune while often matching RLHF's results.
How it works
DPO starts from a supervised-fine-tuned model plus a frozen copy of it
used as the reference policy. Training data is triples of
(prompt, chosen, rejected). For each triple it computes the
log-probability the policy and the reference assign to both responses,
then minimizes a logistic lossLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. that pushes
the chosen response's policy-to-reference ratio above the rejected
one's. A beta hyperparameter controls how far the policy may drift
from the reference.
The derivation shows that under the Bradley-Terry preference model the optimal policy can be written in closed form in terms of the reference, so no explicit reward model and no sampling loop are needed. Each update is an ordinary gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. step on a supervised objective.
When it breaks
- Two models in memory. Policy and reference are loaded together, so a naive setup roughly doubles memory versus plain supervised training.
betais the whole ballgame. Too low and the policy collapses toward short, degenerate answers; too high and preferences barely move the model at all.- It can lower the chosen response's probability. The loss constrains only the margin, so training logs commonly show both chosen and rejected log-probs falling — a real effect that is easy to misread as a bug.
- Off-policy data ages badly. Preference pairs sampled from an older checkpoint no longer reflect what the current model produces, and gains flatten or reverse.
See also: RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward., Fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions.
Learn more: LLMs · Paper: Direct Preference Optimization
Mentioned in
Lessons where this comes up in context.
RLHF (Reinforcement Learning from Human Feedback)
RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward.
Scaling Laws
Scaling laws are empirical relationships between a model's loss and its parameter count, dataset size, and compute budget, used to plan how large to train a new model.