How ChatGPT Was Actually Built
A worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
Every lesson so far in this course covers one piece of the pipeline — pretraining, RLHF, inference, security — as its own topic, in isolation. This is the same material read the other way: as one pipeline, in the order it actually happens, where each stage's output is the next stage's input.
Stage 1: pretraining — learning language, not a task
The pipeline starts with pretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability.: predict the next token, over and over, across hundreds of billions to trillions of tokens of internet-scale text. Nothing here is specific to conversation, question-answering, or being helpful — it's the same self-supervised objective LLMs covers, and it's what produces a base model: something that can fluently continue any text, with no particular tendency to be a helpful assistant rather than, say, autocompleting a list of similar-sounding fake product reviews.
Scaling lawsScaling LawsScaling laws are empirical relationships between a model's loss and its parameter count, dataset size, and compute budget, used to plan how large to train a new model. research establishes that, for a fixed compute budget, model size and training-data size trade off against each other on a predictable curve — which is why labs can plan a training run's size in advance rather than discovering it empirically. Pretraining a model in the GPT-3 class (175 billion parameters) on hundreds of billions of tokens is a run measured in thousands of GPUsGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires. running for weeks to months. The stages that come after it — the ones that actually make the model into a usable chat assistant — are smaller by orders of magnitude: a fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions. and RLHF pass measured in thousands of human-written or human-ranked examples, not hundreds of billions of tokens. Pretraining is where nearly all the compute goes; alignmentAlignmentAlignment is whether a trained model's actual objective and behavior match what its designers intended, split into outer and inner alignment. is where nearly all the product comes from.
Stage 2: alignment — teaching it to be an assistant
A base model completes text; it doesn't reliably follow instructions or refuse harmful requests. AI Safety & Alignment covers this stage's mechanics in depth — supervised fine-tuning on example conversations, then RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward.: collecting human rankings of candidate responses, training a rewardRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one. model on those rankings, and optimizing the base model against it. The InstructGPT paper's own headline finding is the clearest evidence this stage matters as much as it does: human raters preferred outputs from a 1.3-billion-parameter model that had been through this pipeline over outputs from the 175-billion-parameter base model it started from — more than 100x smaller, and still preferred, because the thing raters were judging wasn't raw capability. It was whether the model actually did what was asked.
That preference gap is the entire justification for running an expensive alignment stage at all: pretraining alone produces capability, not an assistant.
Stage 3: inference — turning weights into a running service
An aligned model is still just a file full of weights until it's served. Inference & Serving covers the specific engineering here — the KV cacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable. that avoids recomputing attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. over the whole conversation on every new token, batchingBatching (Continuous Batching)Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish. that shares GPU work across many users' requests at once, quantizationQuantizationQuantization reduces a model's numerical precision (e.g. 16-bit to 4-bit) to shrink memory footprint and speed up inference, at some cost to accuracy. that trades a little precision for meaningfully less memory and cost per request. None of this changes what the model knows or how it was trained — it's entirely about the difference between a model that works in a research notebook and one that can serve a global product at an acceptable cost and latency, which is its own engineering discipline distinct from everything in the two stages before it.
Stage 4: security — the system is now a target
The moment a model is reachable by real users, AI Security stops being theoretical. A production chat assistant is a live target for prompt injectionPrompt InjectionPrompt injection is an attack where text an LLM processes — user input, a retrieved document, or a tool's output — contains instructions that override the application's intended behavior. and jailbreaksJailbreakA jailbreak is a prompt crafted to bypass a model's own safety training and get it to produce output it was tuned to refuse. from day one of deployment, which is why shipping labs run red teamingRed TeamingRed teaming is the practice of deliberately probing a deployed model for harmful, unsafe, or exploitable behavior before and after release. before release and keep running it after — new attack techniques keep surfacing against models that already shipped, the same way new exploits keep surfacing against any other widely used piece of deployed software. Findings from red teaming often loop back into stage 2, as new refusal examples added to future rounds of alignment data — the pipeline isn't strictly linear in practice; security findings feed back upstream.
The shape of the whole thing
Pretraining turns raw text into a capable base model — the large majority of the compute budget, none of the assistant behavior yet.
Alignment (SFT + RLHF) turns that base model into something that follows instructions and refuses what it's meant to refuse — a small fraction of the compute, most of what makes it usable.
InferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. turns trained weights into a service that responds in real time, at the cost and scale a real product needs.
Security treats the now-live service as a target from day one, and feeds what it finds back into the next round of alignment.
Studied one lesson at a time, pretraining, alignment, inference, and security read as four unrelated topics. Studied as this pipeline, they're four stages of the same problem — each one a prerequisite the next stage takes for granted, and each one, on its own, an answer to a completely different question than the others.
Interpretability
Reverse-engineering what a trained network's weights actually compute — probing, superposition, sparse autoencoders, and circuits, instead of judging a model by its outputs alone
Applied & Agentic Systems
How prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application