Red Teaming
Red teaming is the practice of deliberately probing a deployed model for harmful, unsafe, or exploitable behavior before and after release.
Red teaming is adversarial testing borrowed directly from security practice: a team (internal, external, or both) deliberately tries to break a model — eliciting harmful content, finding jailbreaksJailbreakA jailbreak is a prompt crafted to bypass a model's own safety training and get it to produce output it was tuned to refuse., constructing prompt injectionsPrompt InjectionPrompt injection is an attack where text an LLM processes — user input, a retrieved document, or a tool's output — contains instructions that override the application's intended behavior. — before those failures reach real users or get found by an actual attacker.
How it works
Red teaming happens at multiple stages: before release, to catch failures while there's still time to fix them, and continuously after release, since new techniques keep appearing against already-deployed models. Findings feed back into the system in two ways: retraining (adding refusal examples for newly discovered attack patterns to RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward. data) and deployment-side mitigations (classifiers that catch known attack patterns, rate limits, tool permission changes) that don't require a new training run.
Manual red teaming doesn't scale to the space of possible attacks, so labs increasingly automate part of it: using one model to generate large volumes of candidate attacks against another, then having humans review and triage what the automated pass finds — trading coverage for the judgment a person still brings to deciding what actually counts as a failure.
When it breaks
- Coverage is inherently incomplete. Red teaming finds the failures someone thought to look for; it cannot prove the absence of failures no one has tried yet.
- It's a point-in-time signal on a moving target. A model, its system prompt, and its available tools can all change after a red team's findings were reported, silently invalidating the results.
- Automated red teaming inherits its generator's blind spots. An LLM generating attacks is unlikely to discover an entire category of failure that its own training never exposed it to.
See also: JailbreakJailbreakA jailbreak is a prompt crafted to bypass a model's own safety training and get it to produce output it was tuned to refuse., Prompt InjectionPrompt InjectionPrompt injection is an attack where text an LLM processes — user input, a retrieved document, or a tool's output — contains instructions that override the application's intended behavior., Adversarial ExampleAdversarial ExampleAn adversarial example is an input with a small, often imperceptible perturbation deliberately crafted to make a model produce a wrong or attacker-chosen output.
Learn more: AI Security · Evaluation & Benchmarks
Mentioned in
Lessons where this comes up in context.
- AI Safety & AlignmentWhether a model's own objectives match what we actually want, independent of any attacker — specification gaming, outer vs. inner alignment, why RLHF isn't a complete answer, and scalable oversight
- AI SecurityAdversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
Adversarial Example
An adversarial example is an input with a small, often imperceptible perturbation deliberately crafted to make a model produce a wrong or attacker-chosen output.
PyTorch
PyTorch is the dominant deep learning framework in both research and production, providing tensor computation, GPU dispatch, and automatic differentiation.