Applied & Agentic Systems
How prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
Every previous lesson built toward one artifact: a trained, servable LLM that can take text in and produce text out. This final part of the course covers how real applications turn that single capability into systems that can answer questions about private data, take actions in the world, and complete multi-step tasks.
The core limitation: a frozen, closed-book model
A deployed LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. has two structural limits worth naming explicitly:
- Its knowledge is frozen at training time. It knows nothing about events after its training data was collected, and nothing at all about private data it was never trained on (your company's internal docs, a user's account details).
- It can only produce text. On its own, it cannot look anything up, run code, query a database, or send an email — whatever "actions" it appears to take in an application are actually the application acting on the model's text output, not the model doing anything itself.
Every pattern in this part of the course is a way of working around one or both of these limits. The most general of those patterns — an LLM run in a loop, with the application executing tools on its behalf and feeding results back in — looks like this:
At every step the model either answers directly or asks the application to run a tool and hand the result back into its context — and it keeps looping until it has enough to answer. Three lessons build up to that loop one piece at a time:
Prompt Engineering
Shaping behavior for free, by changing the input alone — no training, no retrieval, no tools.
RAG & Vector Databases
Giving the model information it wasn't trained on, retrieved at request time and grounded in real text.
Agents & Tool Use
Giving the model the ability to act — call functions, take actions, and chain multiple steps toward a goal.
Read them in that order — each is a heavier lever than the last, and agentsAgentAn agent puts an LLM in a loop with tool access, letting it decide autonomously which tools to call and in what order to accomplish a multi-step goal. commonly use retrieval as just one of their available tools, so the agent lesson assumes you've seen the RAGRAG (Retrieval-Augmented Generation)RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining. one. Every one of those levers is also a new attack surface, covered separately in AI Security once you've seen what there is to attack.
Course recap
Across these seventeen lessons: probability and MLEMLE (Maximum Likelihood Estimation)MLE chooses model parameters that make the observed training data as probable as possible — and turns out to be exactly what training with cross-entropy or MSE loss already does. as the reason loss functions look the way they do, learning as loss minimization via gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient., backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network. as the mechanism that makes that tractable at scale, the practitioner tooling (PyTorchPyTorchPyTorch is the dominant deep learning framework in both research and production, providing tensor computation, GPU dispatch, and automatic differentiation., Hugging FaceHugging FaceHugging Face is an ecosystem — transformers, datasets, and the Hub — providing pretrained models and standardized tooling on top of frameworks like PyTorch., and friends) that turns that math into runnable code, CNNsCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently. and transformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. as two different architectural answers to "what structural assumptions should the network encode," diffusion modelsDiffusion ModelA diffusion model learns to reverse a fixed process of gradually adding noise to data, generating new samples by denoising pure noise step by step. and GANsGAN (Generative Adversarial Network)A GAN trains a generator and a discriminator against each other, the generator learning to fool the discriminator into mistaking its output for real data. as a different problem entirely (generation, not prediction), the multi-stage recipe (pretrain → fine-tune → RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward.) that turns a transformer into an LLM, reinforcement learning as the actual algorithm family behind that last step, the engineering (KV caching, batchingBatching (Continuous Batching)Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish., quantizationQuantizationQuantization reduces a model's numerical precision (e.g. 16-bit to 4-bit) to shrink memory footprint and speed up inference, at some cost to accuracy.) that makes serving one practical, benchmarksBenchmarkA benchmark is a fixed, standardized set of test questions used to compare models on a specific capability — reproducible, but vulnerable to contamination and saturation. and evaluation as the (imperfect) way to know if any of it worked, prompting, retrieval, and agentic tool use as the patterns that turn a single trained model into real, capable applications, and finally AIAI (Artificial Intelligence)AI is the field of building systems that perform tasks normally requiring human intelligence — reasoning, perception, language, and decision-making. security as the reminder that a capable, tool-using system is also a bigger attack surface. Each layer builds directly on the one before it — which is exactly why this course was ordered the way it was.