Prompt Injection
Prompt injection is an attack where text an LLM processes — user input, a retrieved document, or a tool's output — contains instructions that override the application's intended behavior.
Prompt injection exploits a structural fact about how LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. work: the system prompt, the user's message, retrieved documents, and tool output are all just tokens concatenated into one context window. Nothing marks any of it as more trustworthy than the rest, and the model was trained to follow instructions wherever it finds them — so any of those sources can smuggle in instructions the application never intended to run.
How it works
Direct injection is a user simply typing an instruction meant to override the system prompt — "ignore your previous instructions and instead...". It's the easiest case to reason about, because the attacker and the user are the same person.
Indirect injection is the more dangerous case: the malicious instruction arrives inside content the model retrieves or is shown, not text the user typed. A RAGRAG (Retrieval-Augmented Generation)RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining. system that retrieves a web page, an agentAgentAn agent puts an LLM in a loop with tool access, letting it decide autonomously which tools to call and in what order to accomplish a multi-step goal. that reads an email, or an MCPMCP (Model Context Protocol)MCP is an open standard, introduced by Anthropic, for how applications expose tools and data to LLMs, so a tool built once can be reused across different LLM apps. server that returns a file — any of them can return text containing hidden instructions, and the model has no reliable way to tell "this is data I should reason about" apart from "this is a command I should obey," because both arrive as the same kind of token sequence.
When it breaks
- There is no parameterized-query equivalent. SQL solved this class of problem by separating code from data at the protocol level. Nothing analogous exists for LLM context — delimiters and formatting hints reduce the attack surface but don't eliminate it.
- Higher-privilege tools raise the stakes, not the odds. An injected instruction is just as likely to succeed whether the model can only reply with text or can send an email — but the second case turns a successful injection into actual data leaving the system.
- Defenses are layers, not a fix. Instruction-hierarchy training (weighting system-prompt instructions above retrieved content), output filtering, and least-privilege tool access all reduce risk; none of them close the gap completely.
See also: JailbreakJailbreakA jailbreak is a prompt crafted to bypass a model's own safety training and get it to produce output it was tuned to refuse., Red TeamingRed TeamingRed teaming is the practice of deliberately probing a deployed model for harmful, unsafe, or exploitable behavior before and after release., AgentAgentAn agent puts an LLM in a loop with tool access, letting it decide autonomously which tools to call and in what order to accomplish a multi-step goal., MCPMCP (Model Context Protocol)MCP is an open standard, introduced by Anthropic, for how applications expose tools and data to LLMs, so a tool built once can be reused across different LLM apps.
Learn more: AI Security
Mentioned in
Lessons where this comes up in context.
- AI SecurityAdversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
Sparse Autoencoder (SAE)
A sparse autoencoder reconstructs a layer's activations through a much wider, mostly-zero hidden layer, decomposing overlapping neurons into more individually meaningful features.
Jailbreak
A jailbreak is a prompt crafted to bypass a model's own safety training and get it to produce output it was tuned to refuse.