Fine-tuning
Fine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions.
Fine-tuning takes a pretrainedPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability. base model
and continues training it on a smaller, curated dataset shaped like the
target behavior — e.g. (instruction, ideal response) pairs for
supervised fine-tuning (SFT). It's far cheaper than pretraining since
the model already has language and knowledge; fine-tuning mostly teaches
behavior, not new facts. PEFT methods like LoRA fine-tune by
training a small number of additional parameters instead of the whole
model, making fine-tuning practical even for very large models.
How it works
Supervised fine-tuning is ordinary next-token training on a new corpus, with two differences. First, examples are formatted into a chat template with role markers, and the loss is usually masked so only the assistant's tokens contribute — the model is graded on its answers, not on reproducing the prompt. Second, the learning rate is one to two orders of magnitude lower, since the goal is to nudge existing representations rather than build them.
Typical runs are a few epochs over thousands to tens of thousands of examples. Full fine-tuning updates every weight and carries optimizer state for each; LoRALoRA (Low-Rank Adaptation)LoRA is a parameter-efficient fine-tuning method that trains a small number of additional low-rank parameters instead of updating an entire model's weights. and other PEFT methods train a small adapter and keep the base weights frozen.
When it breaks
- Catastrophic forgetting. Training hard on a narrow domain degrades unrelated abilities the base model had. Mixing in general data or lowering the learning rate mitigates it.
- Small datasets overfit fast. With a few thousand examples the model memorizes phrasing within an epoch or two — hold out a real validation splitTrain/Validation/Test SplitSplitting data into training, validation, and test sets lets you measure a model's ability to generalize to new data instead of just memorizing what it trained on. and watch it rather than training for a fixed step count.
- It teaches style, not facts. Fine-tuning on documents rarely makes a model recall them reliably, and can raise hallucinationHallucinationA hallucination is a confidently stated, fluent LLM output that is factually wrong, a direct consequence of models being trained to produce plausible text rather than verified facts. by training it to answer confidently about things it doesn't know.
- Template drift. Serving with a chat template different from the training one quietly costs much of the gain.
See also: PretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability., RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward., LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama.
Learn more: LLMs · Tooling & The Dev Stack
Mentioned in
Lessons where this comes up in context.
- AI SecurityAdversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
- Computer VisionConvolutions, pooling, CNNs, transfer learning
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- Optimization & Training DynamicsWhy plain gradient descent isn't what actually trains modern models — Adam's mechanics, why warmup exists, how learning-rate schedules are shaped, and what batch size actually trades off at scale
- Prompt EngineeringShaping an LLM's output by changing the input alone — few-shot examples, chain-of-thought, and system prompts — with no training involved
- RAG & Vector DatabasesRetrieval-augmented generation grounds an LLM's answers in retrieved text at request time — and the vector databases and similarity search that make it work
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
Pretraining
Pretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability.
LoRA (Low-Rank Adaptation)
LoRA is a parameter-efficient fine-tuning method that trains a small number of additional low-rank parameters instead of updating an entire model's weights.