LLM (Large Language Model)
An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama.
A large language model (LLM) is a large transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. trained in a multi-stage recipe: pretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability. on huge amounts of text to predict the next token, then fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions. (often with RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward.) to follow instructions and match human preferences. Examples include GPT, Claude, Gemini, and Llama. Unlike earlier AI systems built one-per-task, a single LLM can be prompted into translation, coding, summarization, and many other tasks without task-specific training.
How it works
At inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. an LLM runs a loop. The prompt is tokenized and pushed through the whole stack once — the prefill pass, which is compute-bound and fully parallel. Generation then proceeds one token at a time: the model emits logits over the vocabulary, a decoding strategyDecoding Strategies (Temperature, Top-k, Top-p)Decoding strategies turn an LLM's next-token probability distribution into an actual chosen token, using temperature, top-k, or top-p (nucleus) sampling. selects a token, and that token is appended and fed back in. Every step reuses the cached keys and values of all previous positions via the KV cacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable., so per-step cost stays roughly constant instead of re-reading the entire context.
The model holds no state between requests. Conversation history, retrieved documents, and system instructions are all simply tokens prepended to the context on every call.
When it breaks
- Context limits are hard. Past the trained window, output degrades or the request is rejected. Long contexts also show "lost in the middle" behavior, where facts placed mid-prompt are used less reliably than those at either edge.
- Decoding is memory-bandwidth-bound. Throughput is limited by streaming weights and cache out of GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires. memory, not by arithmetic — which is why larger batches raise tokens per second per dollar but not per request.
- Confident errors. Fluent output carries no calibration, so hallucinationHallucinationA hallucination is a confidently stated, fluent LLM output that is factually wrong, a direct consequence of models being trained to produce plausible text rather than verified facts. is the default failure mode rather than an error message.
- Nondeterminism. Identical prompts can produce different text across batch sizes even at temperature zero.
See also: TransformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs., PretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability., TokenizationTokenizationTokenization converts raw text into a sequence of integers a model can process, typically via subword schemes like byte-pair encoding (BPE)., Fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions.
Learn more: LLMs · Wikipedia: Large language model
Mentioned in
Lessons where this comes up in context.
- Agents & Tool UseGiving an LLM the ability to take actions and chain multiple steps together — tool use, MCP, and the agent loop
- AI SecurityAdversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
- Applied & Agentic SystemsHow prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
- Attention & TransformersThe core architecture behind modern AI
- Evaluation & BenchmarksHow models are actually scored — benchmark suites, leaderboards, eval methodology, and why a high benchmark score doesn't guarantee good real-world performance
- Generative ModelsGANs, VAEs, and diffusion models — how AI generates new images, audio, and video, as opposed to classifying or understanding existing content
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- Inference & ServingBatching, KV caching, quantization, latency/throughput tradeoffs
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- ML FundamentalsSupervised/unsupervised learning, loss functions, gradient descent
- Multimodal ModelsHow a single model handles more than one kind of input — CLIP's shared embedding space, how vision gets fed into a language model as tokens, and what actually breaks
- Neural Networks & BackpropFrom Karpathy's micrograd approach — building a tiny neural net and stepping through forward/backward passes
- Optimization & Training DynamicsWhy plain gradient descent isn't what actually trains modern models — Adam's mechanics, why warmup exists, how learning-rate schedules are shaped, and what batch size actually trades off at scale
- Probability & Statistics FoundationsDistributions, Bayes' theorem, and maximum likelihood estimation — the math that loss functions and uncertainty in ML are actually built on
- Prompt EngineeringShaping an LLM's output by changing the input alone — few-shot examples, chain-of-thought, and system prompts — with no training involved
- RAG & Vector DatabasesRetrieval-augmented generation grounds an LLM's answers in retrieved text at request time — and the vector databases and similarity search that make it work
- Reinforcement LearningMDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
RNN (Recurrent Neural Network)
An RNN is a neural network for sequences that processes input one step at a time, carrying a hidden state forward — the main predecessor to transformers.
Tokenization
Tokenization converts raw text into a sequence of integers a model can process, typically via subword schemes like byte-pair encoding (BPE).