Language Models

LLMs

Tokenization, embeddings, pretraining vs fine-tuning, RLHF basics

The transformer is an architecture. A large language model is what you get by taking that architecture, scaling it up, and training it in a specific multi-stage recipe on text. This lesson walks through that recipe end to end.

The whole pipeline, before we break it into pieces:

Each arrow is the same transformer being trained again — what changes from stage to stage is the data and the objective, never the architecture.

Tokenization: turning text into numbers

Neural networks operate on numbers, not characters. Tokenization is the step that converts raw text into a sequence of integers the model can consume. The simplest options are flawed:

  • Word-level: one token per word — but "run," "runs," "running," and "runner" all become unrelated tokens, and any word not seen during training (a typo, a new product name) has no representation at all.
  • Character-level: no unknown-word problem, but sequences get very long (a sentence becomes dozens of tokens instead of a handful), which is expensive since attention's cost grows with sequence length.

Modern LLMs use subword tokenization (commonly byte-pair encoding, BPE) as a middle ground: common words get their own single token ("the," "is"), while rarer or unfamiliar words get split into smaller frequent pieces ("tokenization" might become "token" + "ization"). This keeps sequences reasonably short while still being able to represent any string, including words never seen during training, by falling back to smaller and smaller pieces.

Try it on real text below — this runs an actual BPE tokenizer (the same encoding used by GPT-3.5/4) live in your browser, no server involved:

Each token ID is then converted into an embedding — a learned vector of numbers (typically hundreds to thousands of dimensions) that represents that token's meaning in a way the network can compute with. Tokens used in similar contexts during training end up with similar embedding vectors — this is learned entirely from data, not hand-specified, and is the raw input that the transformer's attention layers then operate on.

BPE isn't the only subword scheme in use — the alternatives differ in how they decide which pieces belong in the vocabulary:

SchemeHow pieces are chosenUsed by
BPE (byte-pair encoding)Greedily merge the most frequent adjacent pairGPT family, Llama
WordPieceMerge whichever pair most increases the training data's likelihood, not just frequencyBERT
UnigramStart from a large candidate vocabulary and prune the pieces that hurt likelihood leastT5, ALBERT

SentencePiece isn't a fourth scheme — it's a language-agnostic implementation that runs BPE or Unigram as a library, treating whitespace as an ordinary character rather than a word boundary. That matters for languages like Japanese or Chinese, which don't use spaces to separate words the way English does.

Pretraining: next-token prediction at scale

Pretraining is self-supervised learning (see ML Fundamentals) applied to a massive corpus of text: feed the model a sequence, hide the next token, have it predict a probability distribution over the entire vocabulary for what comes next, and use cross-entropy loss against the actual next token. Repeat this across trillions of tokens of text (web pages, books, code, and more).

Nobody labels this data by hand — the "label" at every position is simply whatever word actually comes next in real text. This is why pretraining can use practically the whole internet as training data, and it's the single largest source of everything an LLM "knows": grammar, facts, reasoning patterns, coding conventions, style — all absorbed purely from learning to predict what word comes next, extremely well, across a huge and varied corpus.

The result of pretraining alone is called a base model: genuinely capable of producing fluent, often accurate text, but not naturally inclined to follow instructions or hold a helpful conversation — a base model trained to predict "what comes next" on internet text is just as likely to continue a question with more questions (that's a common pattern in its training data) as it is to answer it.

How much memory do the weights alone need?

Memory for parameters is just a multiplication: bytes = parameters × bytes per parameter. At fp32 that's 4 bytes each, fp16/bf16 2, int8 1, int4 half a byte.

Parametersfp16int8int4
7B~14 GB~7 GB~3.5 GB
70B~140 GB~70 GB~35 GB

So a 7B model at fp16 fits on a single 24 GB consumer GPU, while a 70B model at the same precision needs roughly two 80 GB datacenter GPUs just to hold the weights — before any activations or KV cache. Quantizing to int4 is what drops that same 70B model to ~35 GB and onto one card.

Training is far worse than this table suggests: the optimizer state alone (Adam keeps two extra values per parameter) plus gradients typically pushes the requirement to several times the inference footprint.

Decoding: turning probabilities into text

At each generation step (see Attention & Transformers on causal masking), the model doesn't output a single next token — it outputs a full probability distribution over the entire vocabulary. Decoding is the strategy for turning that distribution into an actual chosen token, and it has a real effect on output quality and character:

  • Greedy decoding: always pick the single highest-probability token. Deterministic, but tends to produce repetitive, bland text — once a slightly-suboptimal token is picked, the model has no way to reconsider.
  • Temperature: before sampling, rescale the probability distribution — low temperature (below 1) sharpens it toward the most likely tokens (closer to greedy, more focused/deterministic); high temperature (above
    1. flattens it (more random, more varied, higher risk of incoherence).
  • Top-k sampling: restrict sampling to only the k most likely next tokens, discarding the long tail of unlikely ones before sampling.
  • Top-p (nucleus) sampling: instead of a fixed count, keep the smallest set of tokens whose cumulative probability exceeds p (e.g. 0.9) — adapts the candidate pool size to how confident the model is at each step, unlike top-k's fixed cutoff.

Production chat products typically combine temperature with top-p or top-k rather than using any one alone, tuned to balance coherence against repetitiveness for the specific product.

Fine-tuning: turning a base model into an assistant

Fine-tuning takes a pretrained base model and continues training it on a smaller, curated dataset shaped like the behavior you actually want. The most common form for LLMs is supervised fine-tuning (SFT): train on examples of (instruction, ideal response) pairs, written or curated by humans, so the model learns the pattern of being a helpful assistant rather than merely continuing text.

This is dramatically cheaper than pretraining — you need thousands to low-millions of high-quality examples, not trillions of tokens — because the model already learned language and world knowledge during pretraining; fine-tuning mostly teaches it a new behavior, not new knowledge.

RLHF: aligning with human preference, not just imitation

SFT alone has a ceiling: it only teaches a model to imitate the specific example responses it was shown, and writing an "ideal" response for every possible situation doesn't scale. RLHF (reinforcement learning from human feedback) addresses this differently:

  1. Collect preference data: show human raters multiple model responses to the same prompt, and have them rank which is better.
  2. Train a reward model: a separate model learns to predict, from those rankings, a numeric score for "how good is this response," generalizing beyond the exact examples raters saw.
  3. Optimize the LLM against that reward: using reinforcement learning, adjust the LLM's parameters to produce responses that score highly according to the reward model — in effect, using human preferences (via the learned reward model as a proxy) as the training signal, instead of a single "correct" answer.

The intuition for why this helps beyond SFT: it's often much easier for a human to judge which of two responses is better than to write the ideal response from scratch, especially for open-ended tasks. RLHF lets that easier judgment signal shape the model at scale, via the reward model acting as a stand-in for a human rater on every training example.

A newer, simpler alternative called DPO (Direct Preference Optimization) achieves a similar effect without the separate reward-model-training and reinforcement-learning steps — it optimizes the LLM directly on the preference-ranked pairs using a loss function derived to have the same optimum RLHF would reach. It's become popular because it's simpler to implement and tune while often matching RLHF's results, though RLHF-style pipelines are still widely used, particularly at the largest labs.

How big should a model be? Scaling laws

History & Landscape mentioned that transformers scale predictably — this can be made precise. Scaling laws (most influentially studied in OpenAI's 2020 paper and DeepMind's 2022 "Chinchilla" paper) empirically relate a model's final loss to three quantities: parameter count, dataset size (tokens trained on), and compute budget. The key practical finding from Chinchilla specifically: many earlier large models were under-trained relative to their size — for a fixed compute budget, a smaller model trained on more data often beats a larger model trained on less, contrary to the earlier assumption that bigger parameter counts were the main lever to pull. This directly shaped how labs allocate compute between "make the model bigger" and "train it on more data" when planning a new training run.

Why LLMs hallucinate

A hallucination is a confidently stated, fluent output that's factually wrong. It's a direct consequence of what the model actually is: next-token prediction trained to produce plausible-sounding text, not a database lookup with a built-in notion of "I don't actually know this." When a model doesn't have reliable knowledge about something, the training objective still rewards it for producing some fluent continuation — it has no mechanism to instead output "I don't know" unless that behavior was specifically reinforced during fine-tuning or RLHF.

This isn't a bug that gets patched

Hallucination isn't a glitch to be fixed in the next model version — it's a structural consequence of the training objective (predict plausible next tokens). Better fine-tuning and RLHF can reduce how often it happens, but the only reliable fix for a specific factual question is grounding the answer in retrieved, verifiable text — see RAG, next.

This is precisely the problem RAG is built to reduce: grounding generation in retrieved, verifiable text instead of relying purely on what the model "remembers" from pretraining.

Putting the pipeline together

Tokenize the training corpus — raw text becomes integer token sequences.

Pretrain on next-token prediction, self-supervised, across trillions of tokens. Output: a base model — fluent and knowledgeable, but not naturally instruction-following.

Supervised fine-tune (SFT) on curated (instruction, ideal response) pairs, teaching the pattern of being a helpful assistant.

RLHF (or DPO): optimize against a learned human-preference signal. Output: an instruction-tuned "chat" model — what you actually interact with as a product.

The three stages differ along every axis except the model itself:

StageDataSignalRelative scale
PretrainingRaw text, unlabeledThe actual next tokenTrillions of tokens
SFTCurated instruction/response pairsThe written ideal responseThousands to millions
RLHF / DPORanked response pairsWhich response a human preferredThousands to millions

Each stage is training the same underlying transformer — nothing about the architecture changes between stages, only the data and objective. This is also why the same base model can be fine-tuned differently for different products: a coding assistant and a general chat assistant can share a pretrained base and diverge only at the fine-tuning stage.

Recap and what's next

Tokenization turns text into a form transformers can process; pretraining on next-token prediction is where nearly all of a model's raw knowledge and language ability comes from; fine-tuning and RLHF shape that knowledge into a helpful, instruction-following assistant. RLHF leaned on reinforcement learning without fully explaining it — the next lesson covers that paradigm properly, before returning to the inference and serving side of the stack.

On this page