AI Security
Adversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
Everything in Prompt Engineering, RAG & Vector Databases, and Agents & Tool Use made a deployed LLM more capable — able to answer with current information, act on the world, chain steps toward a goal. Every one of those capabilities is also a new surface an attacker can push on. This lesson covers that side: how a working system gets attacked, not whether it was built with the right goals in mind.
AI security vs. AI safety: two different questions
It's worth separating these precisely, because the terms get used interchangeably and they aren't the same problem:
- AIAI (Artificial Intelligence)AI is the field of building systems that perform tasks normally requiring human intelligence — reasoning, perception, language, and decision-making. security asks: can someone else make this system do something it wasn't supposed to do? The system may be working exactly as designed — the failure is an external attacker finding a way to abuse it, the same category of problem as a SQL injection or a buffer overflow.
- AI safetyAI SafetyAI safety is the question of whether a system's own behavior matches what its designers actually want, independent of any attacker. and alignmentAlignmentAlignment is whether a trained model's actual objective and behavior match what its designers intended, split into outer and inner alignment. asks: does the system's own behavior match what we actually want, even with no attacker involved? A model that confidently states wrong information, or pursues a training objective in a way its designers didn't anticipate, is a safety problem with nobody attacking it at all.
This lesson is entirely about the first question. Everything below is a way of getting a model to do something outside its intended use, by someone other than the people who built the application. The second question — specification gaming, outer vs. inner alignment, and why RLHF alone doesn't settle it — is its own lesson: AI Safety & Alignment.
Why LLMs are unusually hard to secure
Traditional software has a clean way to separate instructions from data: a SQL query's parameters can't be executed as code, a web form's input gets escaped before it's rendered. An LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. has no equivalent boundary. Its context window is one sequence of tokens — the system prompt, the user's message, a retrieved document, a tool's output — and the model was trained to follow instructions wherever they appear in that sequence. Nothing at the architecture level marks one span of tokens as "commands" and another as "data to reason about." That single structural fact is the root cause of most of what follows.
Prompt injection: attacking the application
Prompt injectionPrompt InjectionPrompt injection is an attack where text an LLM processes — user input, a retrieved document, or a tool's output — contains instructions that override the application's intended behavior. is an attacker getting the model to follow instructions the application never intended it to receive. The simplest version is direct injection: a user typing "ignore your previous instructions and instead..." — easy to reason about, because the attacker and the user are the same person, and the blast radius is usually limited to that one user's own session.
Indirect injection is the version that matters for anything with real capability. The malicious instruction doesn't come from the user at all — it arrives inside content the model is asked to process:
An application gives an agent a legitimate goal — "summarize this webpage," "check my inbox for anything urgent."
As part of that goal, the agent retrieves or is shown content it didn't author: a web page, a document via RAGRAG (Retrieval-Augmented Generation)RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining., an email, a result from an MCPMCP (Model Context Protocol)MCP is an open standard, introduced by Anthropic, for how applications expose tools and data to LLMs, so a tool built once can be reused across different LLM apps. tool.
That content contains text written to look like an instruction — hidden in a webpage's invisible text, buried in an email's body, appended to a file's contents.
The model has no reliable way to tell "this is data I was asked to summarize" apart from "this is a command I should obey," because both arrived as the same kind of token sequence — so it follows the injected instruction.
The person who never typed anything malicious is the one whose agent just got hijacked.
Jailbreaks: attacking the model's own training
A jailbreakJailbreakA jailbreak is a prompt crafted to bypass a model's own safety training and get it to produce output it was tuned to refuse. is a different target: instead of smuggling instructions past an application, it's a user, working entirely within their own conversation, trying to get the model to produce output its own safety fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions. was meant to refuse. Common techniques include role-play framing ("play a character with no restrictions"), hypothetical framing ("for a story I'm writing..."), and many-shot jailbreaking — filling a long context window with turns of a fictional dialogue where an assistant complies with similar requests, so the model continues the pattern it's conditioning on.
Jailbreaks and prompt injection are different attack surfaces, but they compound: an indirectly injected instruction can itself be phrased as a jailbreak attempt, using both techniques in the same attack.
Data exfiltration: when a successful attack has somewhere to go
Prompt injection against a model that can only reply with text is contained — worst case, it says something embarrassing. Prompt injection against an agentAgentAn agent puts an LLM in a loop with tool access, letting it decide autonomously which tools to call and in what order to accomplish a multi-step goal. with real tool access is a different order of problem: the same successful attack can now send an email, write a file, or make an API call, turning a hijacked context window into actual data leaving the system.
A well-known concrete pattern illustrates why: an agent that can render
markdown can be tricked into outputting an image link whose URL encodes
stolen data as a query parameter — .
The client fetches the "image" to render it, which sends that query
string to the attacker's server, and the secret leaks with no
confirmation dialog and nothing for a user to click. The lesson isn't
the specific trick — it's that any channel an agent can be made to
call out through (an image, a link, an API request) is a potential
exfiltration path once an attacker controls what the agent says.
An agent scoped to read-only search has a bounded worst case: a successful injection can waste a request or return a wrong answer. An agent scoped to send email, write files, and make arbitrary HTTP requests has an effectively unbounded worst case — the same successful injection can now act on that access.
The defense that actually changes the outcome isn't better prompt wording, which only changes how likely an injection is to succeed. It's least-privilege tool access: granting exactly the tools and permissions a task needs, and nothing more, so that even a successful injection has a small blast radius by construction. This is the same principle a security team applies to any service account — the model just happens to be the thing making the request.
Adversarial examples: attacking the model's perception
Adversarial examplesAdversarial ExampleAn adversarial example is an input with a small, often imperceptible perturbation deliberately crafted to make a model produce a wrong or attacker-chosen output. target something more fundamental than instructions — the model's decision boundary itself. Ian GoodfellowIan GoodfellowInvented generative adversarial networks (GANs) in 2014, the architecture that first made realistic image generation possible by pitting two networks against each other. and collaborators showed in 2014 that a small, structured perturbation to an image, imperceptible to a person, could make a classifier confidently output the wrong label. The perturbation isn't random noise; it's derived from the model's own gradients, aimed precisely at whatever decision boundary that specific model learned — the same gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. machinery covered in ML Fundamentals, run against the input instead of the weights.
The idea generalizes past vision: research on LLMs has found short, often gibberish-looking suffixes that, appended to a prompt, reliably push a model past its safety training — an automatically discovered jailbreak rather than a hand-crafted one, found the same way: by following gradients toward whatever breaks the model.
Red teaming: finding these failures on purpose, first
Red teamingRed TeamingRed teaming is the practice of deliberately probing a deployed model for harmful, unsafe, or exploitable behavior before and after release. is deliberately probing a model for exactly these failures — before release, so there's still time to fix them, and continuously after release, since new techniques keep appearing against models that already shipped. Findings feed back two ways: retraining (adding refusal examples to RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward. data for newly found patterns) and deployment-side mitigations (classifiers, rate limits, narrower tool permissions) that don't require a new training run at all. Manual red teaming doesn't scale to the space of possible attacks, so labs increasingly automate part of it — using one model to generate candidate attacks against another at volume, with humans reviewing what the automated pass surfaces.
The taxonomy so far
| Attack | Targets | Attacker is | Typical defense |
|---|---|---|---|
| Prompt injection | The application | Anyone who controls data the model reads | Least-privilege tools, instruction hierarchy (weighting system prompts above retrieved content) |
| Jailbreak | The model's own training | The user, in their own conversation | Safety fine-tuning, red teaming |
| Adversarial example | The model's decision boundary | Anyone who can query the model | Adversarial training, input validation |
Red teaming isn't a fourth row — it's the practice of actively looking for all three before someone else finds them first.
Defense in depth, not a patch
None of this is a reason to avoid giving models real capability — RAG and tool use are what make an LLM useful for anything beyond a single Q&A exchange. It's a reason to design the permissions around that capability as carefully as the capability itself, on the assumption that some fraction of what the model reads will eventually be adversarial.
Agents & Tool Use
Giving an LLM the ability to take actions and chain multiple steps together — tool use, MCP, and the agent loop
AI Safety & Alignment
Whether a model's own objectives match what we actually want, independent of any attacker — specification gaming, outer vs. inner alignment, why RLHF isn't a complete answer, and scalable oversight