LLMs
Tokenization, embeddings, pretraining vs fine-tuning, RLHF basics
The transformer is an architecture. A large language modelLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. is what you get by taking that architecture, scaling it up, and training it in a specific multi-stage recipe on text. This lesson walks through that recipe end to end.
The whole pipeline, before we break it into pieces:
Each arrow is the same transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. being trained again — what changes from stage to stage is the data and the objective, never the architecture.
Tokenization: turning text into numbers
Neural networksNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. operate on numbers, not characters. TokenizationTokenizationTokenization converts raw text into a sequence of integers a model can process, typically via subword schemes like byte-pair encoding (BPE). is the step that converts raw text into a sequence of integers the model can consume. The simplest options are flawed:
- Word-level: one token per word — but "run," "runs," "running," and "runner" all become unrelated tokens, and any word not seen during training (a typo, a new product name) has no representation at all.
- Character-level: no unknown-word problem, but sequences get very long (a sentence becomes dozens of tokens instead of a handful), which is expensive since attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer.'s cost grows with sequence length.
Modern LLMs use subword tokenization (commonly byte-pair encoding, BPE) as a middle ground: common words get their own single token ("the," "is"), while rarer or unfamiliar words get split into smaller frequent pieces ("tokenization" might become "token" + "ization"). This keeps sequences reasonably short while still being able to represent any string, including words never seen during training, by falling back to smaller and smaller pieces.
Try it on real text below — this runs an actual BPE tokenizer (the same encoding used by GPT-3.5/4) live in your browser, no server involved:
Each token ID is then converted into an embeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space. — a learned vector of numbers (typically hundreds to thousands of dimensions) that represents that token's meaning in a way the network can compute with. Tokens used in similar contexts during training end up with similar embedding vectors — this is learned entirely from data, not hand-specified, and is the raw input that the transformer's attention layers then operate on.
BPE is trained, not designed. Start with a base vocabulary of individual bytes, so every possible string is representable. Then repeat: count every adjacent pair of symbols in the training corpus, find the most frequent pair , and add a new symbol to the vocabulary, replacing every occurrence. Stop when the vocabulary reaches a target size — typically to for modern models.
Frequent sequences get merged early and end up as single tokens; rare ones
never get merged and stay as several pieces. That is why token counts do
not track word counts. A common English word is usually one token, but a
rare word, a proper noun, a chemical name, or a long identifier in code may
cost three or four. Whitespace is generally merged into the following
token, so the and the are different tokens.
Two consequences that bite in practice. First, non-English text — and especially scripts far from the corpus's distribution — often costs noticeably more tokens per unit of meaning, which means more compute and more money for the same sentence. Second, because tokens are opaque chunks rather than letters, tasks defined over characters (counting letters, reversing a string, rhyming) are genuinely harder for the model than they look — it is not seeing the letters you are seeing.
BPE isn't the only subword scheme in use — the alternatives differ in how they decide which pieces belong in the vocabulary:
| Scheme | How pieces are chosen | Used by |
|---|---|---|
| BPE (byte-pair encoding) | Greedily merge the most frequent adjacent pair | GPT family, Llama |
| WordPiece | Merge whichever pair most increases the training data's likelihood, not just frequency | BERT |
| Unigram | Start from a large candidate vocabulary and prune the pieces that hurt likelihood least | T5, ALBERT |
SentencePiece isn't a fourth scheme — it's a language-agnostic implementation that runs BPE or Unigram as a library, treating whitespace as an ordinary character rather than a word boundary. That matters for languages like Japanese or Chinese, which don't use spaces to separate words the way English does.
Pretraining: next-token prediction at scale
PretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability. is self-supervised learning (see ML Fundamentals) applied to a massive corpus of text: feed the model a sequence, hide the next token, have it predict a probability distributionProbability DistributionA probability distribution assigns a likelihood to each possible value of a random variable — the object every loss function is secretly built to measure the fit of. over the entire vocabulary for what comes next, and use cross-entropy loss against the actual next token. Repeat this across trillions of tokens of text (web pages, books, code, and more).
Nobody labels this data by hand — the "label" at every position is simply whatever word actually comes next in real text. This is why pretraining can use practically the whole internet as training data, and it's the single largest source of everything an LLM "knows": grammar, facts, reasoning patterns, coding conventions, style — all absorbed purely from learning to predict what word comes next, extremely well, across a huge and varied corpus.
The result of pretraining alone is called a base model: genuinely capable of producing fluent, often accurate text, but not naturally inclined to follow instructions or hold a helpful conversation — a base model trained to predict "what comes next" on internet text is just as likely to continue a question with more questions (that's a common pattern in its training data) as it is to answer it.
A language model assigns a probability to a whole sequence of tokens by factoring it into conditionals — each token given everything before it:
This factorization is exact, not an approximation — it is just the chain rule of probability. What the model supplies is the parameterized conditional , produced by a softmax over the vocabulary at each position.
Taking the negative log turns the product into a sum, which is exactly the cross-entropy loss from ML Fundamentals applied at every position:
So "train a language model" and "do classification over the vocabulary at every token position" are the same sentence.
Perplexity is this loss exponentiated:
It is readable as an effective branching factor: a perplexity of 20 means the model is, on average, as uncertain as if it were choosing uniformly among 20 tokens. A uniform model over a vocabulary of size has perplexity exactly , and a perfect model has perplexity 1. Because it depends on the tokenizer — the same text split into more tokens changes the per-token average — perplexity is only comparable across models that share a vocabulary.
Memory for parameters is just a multiplication: bytes = parameters ×
bytes per parameter. At fp32 that's 4 bytes each, fp16/bf16 2,
int8 1, int4 half a byte.
| Parameters | fp16 | int8 | int4 |
|---|---|---|---|
| 7B | ~14 GB | ~7 GB | ~3.5 GB |
| 70B | ~140 GB | ~70 GB | ~35 GB |
So a 7B model at fp16 fits on a single 24 GB consumer GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires., while a 70B
model at the same precision needs roughly two 80 GB datacenter GPUs just to
hold the weights — before any activations or KV cacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable.. Quantizing to
int4 is what drops that same 70B model to ~35 GB and onto one card.
Training is far worse than this table suggests: the optimizer state alone (Adam keeps two extra values per parameter) plus gradients typically pushes the requirement to several times the inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. footprint.
Decoding: turning probabilities into text
At each generation step (see Attention & Transformers on causal masking), the model doesn't output a single next token — it outputs a full probability distribution over the entire vocabulary. Decoding is the strategy for turning that distribution into an actual chosen token, and it has a real effect on output quality and character:
- Greedy decoding: always pick the single highest-probability token. Deterministic, but tends to produce repetitive, bland text — once a slightly-suboptimal token is picked, the model has no way to reconsider.
- Temperature: before sampling, rescale the probability distribution —
low temperature (below 1) sharpens it toward the most likely tokens
(closer to greedy, more focused/deterministic); high temperature (above
- flattens it (more random, more varied, higher risk of incoherence).
- Top-k sampling: restrict sampling to only the k most likely next tokens, discarding the long tail of unlikely ones before sampling.
- Top-p (nucleus) sampling: instead of a fixed count, keep the smallest set of tokens whose cumulative probability exceeds p (e.g. 0.9) — adapts the candidate pool size to how confident the model is at each step, unlike top-k's fixed cutoff.
Production chat products typically combine temperature with top-p or top-k rather than using any one alone, tuned to balance coherence against repetitiveness for the specific product.
The model's raw output at a position is a vector of logits , one score per vocabulary entry. Softmax turns it into a distribution:
where is the temperature. Dividing the logits before exponentiating is the entire mechanism. As the largest logit dominates and the distribution collapses onto — that is greedy decoding, so temperature 0 and greedy are the same thing. At you sample from the model's own distribution unchanged. As grows the exponents shrink toward each other and the distribution flattens toward uniform.
Top-k truncates before sampling: keep the set of indices with the largest logits, zero the rest, renormalize over . The pool size is fixed regardless of how confident the model is.
Top-p, or nucleus sampling, makes that pool adaptive. Sort tokens by probability descending and take the smallest prefix satisfying:
then renormalize over . When the model is confident, a couple of tokens already exceed and the pool is tiny; when it is genuinely uncertain, the pool widens automatically. That adaptivity is the argument for top-p over top-k — a fixed is simultaneously too permissive at confident positions and too restrictive at uncertain ones.
When combined, the conventional order is: scale by temperature, apply the top-k and top-p filters, renormalize, then sample.
Fine-tuning: turning a base model into an assistant
Fine-tuning takes a pretrained base model and continues training it on
a smaller, curated dataset shaped like the behavior you actually want. The
most common form for LLMs is supervised fine-tuning (SFT): train on
examples of (instruction, ideal response) pairs, written or curated by
humans, so the model learns the pattern of being a helpful assistant
rather than merely continuing text.
This is dramatically cheaper than pretraining — you need thousands to low-millions of high-quality examples, not trillions of tokens — because the model already learned language and world knowledge during pretraining; fine-tuning mostly teaches it a new behavior, not new knowledge.
RLHF: aligning with human preference, not just imitation
SFT alone has a ceiling: it only teaches a model to imitate the specific example responses it was shown, and writing an "ideal" response for every possible situation doesn't scale. RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward. (reinforcement learning from human feedback) addresses this differently:
- Collect preference data: show human raters multiple model responses to the same prompt, and have them rank which is better.
- Train a rewardRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one. model: a separate model learns to predict, from those rankings, a numeric score for "how good is this response," generalizing beyond the exact examples raters saw.
- Optimize the LLM against that reward: using reinforcement learning, adjust the LLM's parameters to produce responses that score highly according to the reward model — in effect, using human preferences (via the learned reward model as a proxy) as the training signal, instead of a single "correct" answer.
The intuition for why this helps beyond SFT: it's often much easier for a human to judge which of two responses is better than to write the ideal response from scratch, especially for open-ended tasks. RLHF lets that easier judgment signal shape the model at scale, via the reward model acting as a stand-in for a human rater on every training example.
A newer, simpler alternative called DPODPO (Direct Preference Optimization)DPO aligns an LLM to human preferences directly from ranked response pairs, without the separate reward model and reinforcement learning loop RLHF requires. (Direct Preference Optimization) achieves a similar effect without the separate reward-model-training and reinforcement-learning steps — it optimizes the LLM directly on the preference-ranked pairs using a loss functionLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. derived to have the same optimum RLHF would reach. It's become popular because it's simpler to implement and tune while often matching RLHF's results, though RLHF-style pipelines are still widely used, particularly at the largest labs.
How big should a model be? Scaling laws
History & Landscape mentioned that transformers scale predictably — this can be made precise. Scaling laws (most influentially studied in OpenAI's 2020 paper and DeepMind's 2022 "Chinchilla" paper) empirically relate a model's final loss to three quantities: parameter count, dataset size (tokens trained on), and compute budget. The key practical finding from Chinchilla specifically: many earlier large models were under-trained relative to their size — for a fixed compute budget, a smaller model trained on more data often beats a larger model trained on less, contrary to the earlier assumption that bigger parameter counts were the main lever to pull. This directly shaped how labs allocate compute between "make the model bigger" and "train it on more data" when planning a new training run.
Why LLMs hallucinate
A hallucinationHallucinationA hallucination is a confidently stated, fluent LLM output that is factually wrong, a direct consequence of models being trained to produce plausible text rather than verified facts. is a confidently stated, fluent output that's factually wrong. It's a direct consequence of what the model actually is: next-token prediction trained to produce plausible-sounding text, not a database lookup with a built-in notion of "I don't actually know this." When a model doesn't have reliable knowledge about something, the training objective still rewards it for producing some fluent continuation — it has no mechanism to instead output "I don't know" unless that behavior was specifically reinforced during fine-tuning or RLHF.
This isn't a bug that gets patched
Hallucination isn't a glitch to be fixed in the next model version — it's a structural consequence of the training objective (predict plausible next tokens). Better fine-tuning and RLHF can reduce how often it happens, but the only reliable fix for a specific factual question is grounding the answer in retrieved, verifiable text — see RAGRAG (Retrieval-Augmented Generation)RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining., next.
This is precisely the problem RAG is built to reduce: grounding generation in retrieved, verifiable text instead of relying purely on what the model "remembers" from pretraining.
Putting the pipeline together
Tokenize the training corpus — raw text becomes integer token sequences.
Pretrain on next-token prediction, self-supervised, across trillions of tokens. Output: a base model — fluent and knowledgeable, but not naturally instruction-following.
Supervised fine-tune (SFT) on curated (instruction, ideal response)
pairs, teaching the pattern of being a helpful assistant.
RLHF (or DPO): optimize against a learned human-preference signal. Output: an instruction-tuned "chat" model — what you actually interact with as a product.
The three stages differ along every axis except the model itself:
| Stage | Data | Signal | Relative scale |
|---|---|---|---|
| Pretraining | Raw text, unlabeled | The actual next token | Trillions of tokens |
| SFT | Curated instruction/response pairs | The written ideal response | Thousands to millions |
| RLHF / DPO | Ranked response pairs | Which response a human preferred | Thousands to millions |
Each stage is training the same underlying transformer — nothing about the architecture changes between stages, only the data and objective. This is also why the same base model can be fine-tuned differently for different products: a coding assistant and a general chat assistant can share a pretrained base and diverge only at the fine-tuning stage.
Recap and what's next
Tokenization turns text into a form transformers can process; pretraining on next-token prediction is where nearly all of a model's raw knowledge and language ability comes from; fine-tuning and RLHF shape that knowledge into a helpful, instruction-following assistant. RLHF leaned on reinforcement learning without fully explaining it — the next lesson covers that paradigm properly, before returning to the inference and serving side of the stack.
Multimodal Models
How a single model handles more than one kind of input — CLIP's shared embedding space, how vision gets fed into a language model as tokens, and what actually breaks
Reinforcement Learning
MDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF