Pretraining
Pretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability.
Pretraining is self-supervised training (no human-labeled data needed) of an LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. on a massive text corpus: predict the next token at every position, using the actual next word in real text as the label. This is the primary source of a model's knowledge, grammar, and reasoning patterns. The result is a base model — fluent, but not naturally instruction-following — which is why fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions. and RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward. follow.
How it works
The corpus is tokenizedTokenizationTokenization converts raw text into a sequence of integers a model can process, typically via subword schemes like byte-pair encoding (BPE). once and packed into fixed-length sequences. Each step runs a batch through the model, computes cross-entropy lossLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. between the predicted distribution and the actual next token at every position, and updates the weights — so one sequence of length 4096 supplies 4096 supervised predictions. The labels are just the inputs shifted by one; no annotation is involved anywhere.
Runs last weeks across thousands of accelerators, which requires distributed trainingDistributed Training (FSDP, DeepSpeed, Megatron-LM)Distributed training frameworks split a model's parameters, gradients, and optimizer state across many GPUs so models too large for one GPU can still be trained. to split the model and the batch across devices. The learning rate warms up and then decays on a schedule fixed in advance, and the split between model size and token count is what scaling lawsScaling LawsScaling laws are empirical relationships between a model's loss and its parameter count, dataset size, and compute budget, used to plan how large to train a new model. estimate.
When it breaks
- Data quality dominates. Deduplication and filtering move final quality more than most architecture choices; duplicated documents get memorized and inflate held-out scores.
- Benchmark contamination. Evaluation sets leak into web crawls, so a model can look like it is reasoning when it is reciting.
- Loss spikes. Large runs hit sudden divergence from a bad batch or numerical instability; standard practice is to roll back to a checkpoint and skip the offending data.
- One shot at the schedule. Learning-rate decay assumes a fixed token budget, so you cannot cleanly "train it a bit longer" afterward.
See also: LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama., Fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions., RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward.
Learn more: LLMs · Wikipedia: Self-supervised learning
Mentioned in
Lessons where this comes up in context.
- Attention & TransformersThe core architecture behind modern AI
- Computer VisionConvolutions, pooling, CNNs, transfer learning
- Evaluation & BenchmarksHow models are actually scored — benchmark suites, leaderboards, eval methodology, and why a high benchmark score doesn't guarantee good real-world performance
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- ML FundamentalsSupervised/unsupervised learning, loss functions, gradient descent
Embedding
An embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space.
Fine-tuning
Fine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions.