LoRA (Low-Rank Adaptation)
LoRA is a parameter-efficient fine-tuning method that trains a small number of additional low-rank parameters instead of updating an entire model's weights.
LoRA (Low-Rank Adaptation) is a PEFT (parameter-efficient fine-tuning) method: instead of updating all of a pretrained model's weights during fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions., it freezes the original weights and trains a small number of additional low-rank parameters injected alongside them. This makes fine-tuning large LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. practical on modest hardware, and is why most fine-tuning today doesn't require updating billions of parameters directly.
How it works
Why does LoRA use two small matrices instead of one?
For a frozen weight matrix W of shape (d, k), LoRA learns two small
matrices A of shape (r, k) and B of shape (d, r), and computes
the layer as W x + (alpha / r) * B A x. The rank r is small —
commonly 8 to 64 — so trainable parameters number r * (d + k) instead
of d * k. A is initialized randomly and B to zeros, so the
adapter starts as an exact no-op and training begins from the base
model's behavior unchanged.
Adapters are usually attached to the attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer.
projections, sometimes the feed-forward layers as well. Because the
update is a plain matrix product, B A can be folded back into W
after training, leaving inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. cost
identical to the base model.
When it breaks
- Rank caps capacity. Low
rhandles style and format adaptation well but struggles when the task needs genuinely new knowledge or a new language. alphaandrare coupled. The effective scale isalpha / r, so raisingrwhile holdingalphafixed quietly weakens the update — a frequent cause of "the LoRA did nothing" runs.- Merging interacts with quantization. Adapters trained against a quantizedQuantizationQuantization reduces a model's numerical precision (e.g. 16-bit to 4-bit) to shrink memory footprint and speed up inference, at some cost to accuracy. base can shift behavior when merged back into full-precision weights.
- Coverage matters. Targeting only the query and value projections is the common default, but some tasks need the MLP layers included to move at all.
See also: Fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions., Hugging FaceHugging FaceHugging Face is an ecosystem — transformers, datasets, and the Hub — providing pretrained models and standardized tooling on top of frameworks like PyTorch.
Learn more: LLMs · Tooling & The Dev Stack · Paper: LoRA
Mentioned in
Lessons where this comes up in context.
Fine-tuning
Fine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions.
RLHF (Reinforcement Learning from Human Feedback)
RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward.