Prompt Engineering
Prompt engineering shapes an LLM's output by changing the input phrasing alone — via few-shot examples, chain-of-thought, or system prompts — with no training involved.
Prompt engineering is the cheapest lever for shaping LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. behavior — changing input phrasing with no training involved. Few-shot prompting includes example input/output pairs so the model infers format and style from examples rather than a description. Chain-of-thought prompting asks the model to reason step by step before answering, measurably improving accuracy on multi-step tasks, since each output token can condition on the reasoning already generated. System prompts configure a model's persona and constraints for an entire conversation, separate from user input.
How it works
Nothing about the weights changes. The prompt is just the prefix the model conditions on, and every technique is a way of steering the conditional distribution over the next token toward a better region.
- Few-shot examples pin down format and label space far more reliably than an instruction does, because the model matches the pattern already present in its context.
- Chain-of-thought works because each generated token becomes input for the next, so intermediate steps give later tokens something concrete to condition on — effectively more compute per answer.
- System prompts occupy the earliest positions and are typically given a distinct role in the chat template, which fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions. has taught the model to weight heavily.
Structure matters too: delimiters separating instructions from data, and putting the task after long context, both measurably help.
When it breaks
- Prompts are model-specific and version-fragile. A prompt tuned on one checkpoint can regress on the next release of the same model, with no warning and no diff to inspect.
- Instructions and data are the same channel. Anything retrieved by RAGRAG (Retrieval-Augmented Generation)RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining. or returned by a tool can carry instructions the model follows — the root of prompt injectionPrompt InjectionPrompt injection is an attack where text an LLM processes — user input, a retrieved document, or a tool's output — contains instructions that override the application's intended behavior., which delimiters mitigate but do not solve.
- Stated reasoning is not actual reasoning. Chain-of-thought text is a plausible-sounding narrative, and a confident one can accompany a wrong answer.
- Position effects. Models attend unevenly across long prompts, often neglecting the middle, so burying a constraint there loses it.
See also: LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama., AgentAgentAn agent puts an LLM in a loop with tool access, letting it decide autonomously which tools to call and in what order to accomplish a multi-step goal.
Learn more: Prompt Engineering
Mentioned in
Lessons where this comes up in context.
- Agents & Tool UseGiving an LLM the ability to take actions and chain multiple steps together — tool use, MCP, and the agent loop
- AI SecurityAdversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
- Prompt EngineeringShaping an LLM's output by changing the input alone — few-shot examples, chain-of-thought, and system prompts — with no training involved
- RAG & Vector DatabasesRetrieval-augmented generation grounds an LLM's answers in retrieved text at request time — and the vector databases and similarity search that make it work
Vector Database
A vector database stores embedding vectors and answers nearest-neighbor queries fast, using approximate nearest neighbor (ANN) indexes like HNSW — the backbone of RAG retrieval.
Agent
An agent puts an LLM in a loop with tool access, letting it decide autonomously which tools to call and in what order to accomplish a multi-step goal.