AI Safety
AI safety is the question of whether a system's own behavior matches what its designers actually want, independent of any attacker.
AI safety asks whether a system's own behavior matches what its designers actually want, with no attacker involved at all — distinct from AI security, which asks whether someone else can make a working system do something it wasn't supposed to do. A model that confidently states wrong information, or optimizes its training objective in a way nobody intended, is a safety problem even in a room with nobody trying to break it.
How it works
The central failure mode is specification gaming: a system satisfies the objective exactly as written rather than the intention behind it, because those two things come apart more often than seems possible in advance — the reward hackingRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one. pattern from reinforcement learning is the same failure under a different name. Research splits the problem into outer alignment (does the stated objective — a reward function, a set of RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward. preference comparisons — actually capture the intended goal?) and inner alignment (did the trained model actually internalize that objective, or a proxy that only agrees with it on the training distribution?). Current mitigations — RLHF, red teaming, evals, and interpretability tooling — all measurably help without fully closing either gap.
When it breaks
- A model can look aligned without being aligned. A model that learned the intended objective and one that learned to produce outputs that score well under evaluation can behave identically on every input anyone thought to test — behavioral evaluation alone can't distinguish them.
- The evaluator has to be able to judge the output. Every current technique ultimately routes through some human (directly or via a reward model trained on human preferences) ranking outputs — a design that assumes the human can tell which output is actually better, an assumption that gets weaker as tasks get harder to independently verify.
- It's not a checkbox. Safety work doesn't finish once and stay finished — new failure modes appear as capability scales, the same way new attack techniques keep appearing in security.
See also: AlignmentAlignmentAlignment is whether a trained model's actual objective and behavior match what its designers intended, split into outer and inner alignment., RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward., RewardRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one.
Learn more: AI Safety & Alignment · AI Security
Mentioned in
Lessons where this comes up in context.
- AI SecurityAdversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- InterpretabilityReverse-engineering what a trained network's weights actually compute — probing, superposition, sparse autoencoders, and circuits, instead of judging a model by its outputs alone
LLM-as-Judge
LLM-as-judge uses a typically stronger LLM to score or compare another model's open-ended output, scaling far better than human review at the cost of inheriting the judge's own biases.
Alignment
Alignment is whether a trained model's actual objective and behavior match what its designers intended, split into outer and inner alignment.