Interpretability
Interpretability is the practice of identifying what computation a trained network is actually performing internally, instead of inferring intent purely from its outputs.
Interpretability identifies what a trained network's weights are actually computing internally, rather than inferring intent purely from its inputs and outputs. It's a direct response to a specific gap in behavioral evaluation: a model that has learned the intended objective and one that has learned to merely produce outputs that score well can behave identically on every input anyone thought to test — alignmentAlignmentAlignment is whether a trained model's actual objective and behavior match what its designers intended, split into outer and inner alignment. research increasingly treats checking the mechanism, not just the output, as necessary for closing that gap.
How it works
Probing trains a simple linear classifier on a model's internal activations to test whether some concept is linearly recoverable at a given layer. Individual neurons are frequently polysemantic — responding to several unrelated concepts at once — a consequence of superposition: a network represents far more concepts than it has neurons by encoding them as overlapping directions in activation space rather than one concept per neuron. A sparse autoencoderSparse Autoencoder (SAE)A sparse autoencoder reconstructs a layer's activations through a much wider, mostly-zero hidden layer, decomposing overlapping neurons into more individually meaningful features. is the main current tool for decomposing that superposition back into more individually meaningful features, and a circuit is a set of features and attention heads wired together that implements one specific piece of behavior — the level at which findings become causal claims about how an output was produced, not just correlational ones.
When it breaks
- A compelling explanation isn't automatically a correct one. A found "circuit" can fit cherry-picked examples without holding up under systematic or causal testing — the same failure mode as any post-hoc pattern-finding exercise.
- Probing shows recoverability, not causal use. A concept being linearly decodable from activations doesn't establish that the model's own downstream computation actually reads or acts on it.
- Coverage, not technique, is the bottleneck. Current methods explain specific circuits for specific behaviors in real depth; they don't yet scale to comprehensively auditing an entire model's behavior.
See also: AlignmentAlignmentAlignment is whether a trained model's actual objective and behavior match what its designers intended, split into outer and inner alignment., Sparse AutoencoderSparse Autoencoder (SAE)A sparse autoencoder reconstructs a layer's activations through a much wider, mostly-zero hidden layer, decomposing overlapping neurons into more individually meaningful features., AI SafetyAI SafetyAI safety is the question of whether a system's own behavior matches what its designers actually want, independent of any attacker.
Learn more: Interpretability · AI Safety & Alignment
Mentioned in
Lessons where this comes up in context.
- AI Safety & AlignmentWhether a model's own objectives match what we actually want, independent of any attacker — specification gaming, outer vs. inner alignment, why RLHF isn't a complete answer, and scalable oversight
- InterpretabilityReverse-engineering what a trained network's weights actually compute — probing, superposition, sparse autoencoders, and circuits, instead of judging a model by its outputs alone
Alignment
Alignment is whether a trained model's actual objective and behavior match what its designers intended, split into outer and inner alignment.
Sparse Autoencoder (SAE)
A sparse autoencoder reconstructs a layer's activations through a much wider, mostly-zero hidden layer, decomposing overlapping neurons into more individually meaningful features.