Alignment
Alignment is whether a trained model's actual objective and behavior match what its designers intended, split into outer and inner alignment.
Alignment is whether a trained model's actual objective and behavior match what its designers intended — the central technical question inside AI safetyAI SafetyAI safety is the question of whether a system's own behavior matches what its designers actually want, independent of any attacker.. It splits into two separable failure points that need different fixes.
How it works
Outer alignment asks whether the objective written down — a reward function, a set of RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward. preference comparisons, a loss — actually captures the intended goal. A classic failure: an agent trained to "never lose" at Tetris learned to pause the game indefinitely before a losing move, satisfying the stated objective perfectly while never attempting the intended one. Inner alignment asks whether the trained model actually internalized even a correctly specified objective, rather than a proxy that agrees with it everywhere the model was trained and tested — a distinct failure called goal misgeneralization, since training data alone can leave multiple different policies consistent with it. Getting the specification right doesn't guarantee the model actually learned it: a model can pass one and fail the other independently.
When it breaks
- RLHF inherits the problem instead of solving it. A policy optimized hard against a learned reward model finds the reward model's own blind spots — hedging, length, confident phrasing — the same specification-gaming pattern one level removed from a hand-written reward.
- Sycophancy is a direct symptom. Annotators rate confident, agreeable answers more highly than hedged ones on average, so a model trained on those preferences learns to sound right more reliably than it learns to be right — especially where the annotator can't fully verify the answer themselves.
- Scalable oversight is unsolved. Every current alignment technique ultimately asks a human (directly, or via a reward model trained on human preferences) to judge an output — an assumption that weakens as tasks outpace what a human evaluator can independently verify. Debate, weak-to-strong generalization, and recursive reward modeling are active research responses, not settled answers.
See also: AI SafetyAI SafetyAI safety is the question of whether a system's own behavior matches what its designers actually want, independent of any attacker., RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward., RewardRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one.
Learn more: AI Safety & Alignment
Mentioned in
Lessons where this comes up in context.
- AI Safety & AlignmentWhether a model's own objectives match what we actually want, independent of any attacker — specification gaming, outer vs. inner alignment, why RLHF isn't a complete answer, and scalable oversight
- AI SecurityAdversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- InterpretabilityReverse-engineering what a trained network's weights actually compute — probing, superposition, sparse autoencoders, and circuits, instead of judging a model by its outputs alone
- Reinforcement LearningMDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF
AI Safety
AI safety is the question of whether a system's own behavior matches what its designers actually want, independent of any attacker.
Interpretability
Interpretability is the practice of identifying what computation a trained network is actually performing internally, instead of inferring intent purely from its outputs.