Applied

Agents & Tool Use

Giving an LLM the ability to take actions and chain multiple steps together — tool use, MCP, and the agent loop

RAG & Vector Databases covered giving a model information it wasn't trained on. This lesson covers the other structural limit of a deployed LLM: on its own, it can only produce text — it cannot look anything up, run code, query a database, or send an email. Whatever "actions" it appears to take in an application are actually the application acting on the model's text output, not the model doing anything itself.

Tool use: letting the model take (indirect) actions

Tool use (also called function calling) addresses the can't-do-anything limitation. The model is given a description of available tools — a web search function, a calculator, an internal API — each with a name, a description, and a schema for its inputs. Instead of directly answering, the model can output a structured request to call a specific tool with specific arguments; the surrounding application actually executes that call, and feeds the result back into the model's context so it can use it to continue answering.

Crucially, the model itself never runs code or hits an API — it only ever produces text (here, a structured tool call). The actual execution, authentication, and side effects are entirely the responsibility of the application wrapped around the model. Tool use just gives the surrounding system a clean protocol for extending what the model can effectively accomplish, without changing what the model fundamentally is.

A practical problem tool use created: every application wiring an LLM to its own tools originally had to define its own bespoke integration. MCP (Model Context Protocol) is an open standard, introduced by Anthropic in late 2024, that standardizes how applications expose tools and data to LLMs — letting a tool built once be reused across different LLM-powered applications, the same way a USB standard means a peripheral doesn't need custom wiring for every computer it plugs into.

Agents: chaining tool use over multiple steps

An agent is what you get when you put an LLM in a loop with tool use and let it decide, autonomously, which tools to call and in what order to accomplish a goal — rather than a single request/response with at most one tool call:

Given a goal and the current context, the model decides: respond directly, or call a tool.

If it calls a tool, the application executes it and appends the result to the context.

Repeat from step 1, until the model decides the goal is achieved and produces a final response.

Here's that loop scripted out concretely. Step through it (or watch it run) across three goals — one that needs two tool calls, one that needs none at all, and one where the right move is admitting a search came up empty rather than guessing:

Goal: "Is it warmer in Miami or Seattle right now?"

Click "Run" to watch the agent loop step through this goal.

0 / 8

This is a genuinely different mode from a single Q&A exchange: an agent might search the web, read a file, run a calculation, and check its own intermediate result against the original goal — several LLM calls chained together, each one building on tool results from the previous step, all directed at one overall task rather than one prompt.

The tradeoffs are real: more steps means more opportunities for the model to make a wrong call, and errors can compound across a chain rather than being isolated to one response — which is why practical agent systems usually add guardrails (limits on how many steps to take, confirmation before high-stakes actions, structured output validation at each step) rather than letting a model loop entirely unsupervised.

The cost of a growing context

An agent re-sends its entire context on every iteration — the original goal, plus every tool call and result appended so far. Say the prompt starts at 2,000 tokens and each step adds 1,000 tokens of tool output.

Over 10 steps, the input tokens processed per step are 2k, 3k, 4k … 11k — about 65,000 tokens total, not the 12,000 a single final call would suggest. Run 20 steps and it is roughly 250,000.

The cost grows with the square of the step count, and so does latency, since prefill work scales with prompt length. This is why agent frameworks truncate, summarise, or cache context rather than letting the transcript grow unbounded — and why a step budget is a cost control, not just a safety rail.

Combining agents with retrieval

RAG and tool use aren't mutually exclusive — a well-built agent commonly has retrieval (from RAG & Vector Databases) as one of its available tools ("search internal docs") alongside other tools (web search, code execution, sending a message), and decides which to invoke, and in what sequence, based on the specific request. The pattern underlying all of it is the same one from the start of this lesson: the LLM itself stays a frozen, text-in-text-out model; RAG and tool use are both ways of enriching what goes into that model and acting on what comes out of it.

Recap and what's next

Tool use gives a model a clean protocol for taking indirect action; MCP standardizes how those tools are exposed; agents put the model in a loop so it can chain steps together toward a goal, at the cost of compounding error and growing context. Together with Prompt Engineering and RAG & Vector Databases, this is the full applied toolkit for turning a single trained model into a real, capable application — which the final lesson ties together.

On this page