Inference & Applied Systems

Agent

An agent puts an LLM in a loop with tool access, letting it decide autonomously which tools to call and in what order to accomplish a multi-step goal.

An agent puts an LLM in a loop with access to tools, letting it decide — autonomously, from one step to the next — which tools to call and in what order to accomplish a goal, rather than answering in a single request/response turn. Each step's tool result feeds back into the model's context to inform the next decision, until the model determines the goal is achieved. Practical agent systems add guardrails (step limits, confirmation before high-stakes actions) since errors can compound across a chain of steps.

How it works

The loop is mechanically simple. The application sends the model a system prompt, the conversation so far, and a list of tool schemas — name, description, and a JSON Schema for the arguments. The model replies either with ordinary text or with a structured tool call. The application executes that call itself, appends the result to the message list as a tool role message, and calls the model again. The model never runs anything; it only emits requests. Everything the agent "knows" at step n is whatever is in that growing message list, so the loop is really an exercise in context management: what to append verbatim, what to summarise, and what to drop. Stopping is a policy decision — a final text answer, a done tool, or a hard step cap.

When it breaks

  • Error compounding. Each step conditions on the last, so one bad tool result or misread output steers every later decision. A 95% per-step success rate is about 60% over ten steps.
  • Context exhaustion. Raw tool output — file dumps, HTML, long logs — fills the window fast, and truncating it silently drops the evidence the model needs, which invites hallucination.
  • Loops and thrash. Agents retry a failing call with cosmetic variations instead of changing approach. Step caps and repeated-call detection are load-bearing, not optional.
  • Cost and latency scale with steps. Every step is a full inference pass over a longer prompt, so cost grows worse than linearly.

See also: RAG, MCP, LLM

Learn more: Agents & Tool Use

Mentioned in

Lessons where this comes up in context.

On this page