Tensor Parallelism
Tensor parallelism splits individual weight matrices across multiple GPUs so a model too large for one GPU's memory can still be trained or served.
Tensor parallelism splits individual weight matrices across multiple GPUs, each computing a portion of every matrix multiplication and exchanging partial results — used both to train and to serve models too large to fit on a single GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires.'s memory even after quantizationQuantizationQuantization reduces a model's numerical precision (e.g. 16-bit to 4-bit) to shrink memory footprint and speed up inference, at some cost to accuracy.. It adds inter-GPU communication overhead on the critical path of every forward pass, which is why serving very large models efficiently depends heavily on fast GPU-to-GPU interconnects, not just raw GPU count.
How it works
Inside a transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. block the splits follow
the matrix algebra. A feed-forward layer's first weight matrix is cut
column-wise, so each device produces a slice of the hidden
activations and can apply the
activation functionActivation FunctionAn activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function. locally. The
second matrix is cut row-wise, so each device computes a partial sum
of the output and a single all-reduce adds them together. Attention
splits naturally by head: each device owns a subset of heads and their
KV cacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable., with another all-reduce after the
output projection.
That is two collectives per block, on the critical path, for every token. Degree is therefore kept within one node — typically 2, 4, or 8 GPUs on NVLink — while pipeline or data parallelism spans nodes.
When it breaks
- Communication does not shrink with scale. Each
all-reducemoves activation-sized tensors regardless of how finely the weights are split, so past a point adding GPUs raises latency instead of lowering it. - Interconnect is the deciding factor. The same degree that works over NVLink can be slower than a single device over PCIe, and worse across nodes.
- Divisibility constraints. Head count and hidden size must divide evenly by the degree, which quietly rules out some model/GPU-count combinations.
- Stragglers sync the whole group. Collectives are barriers, so one slow or thermally throttled device sets the pace for every other.
See also: GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires., InferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering.
Learn more: Inference & Serving · Tooling & The Dev Stack
Mentioned in
Lessons where this comes up in context.
Speculative Decoding
Speculative decoding uses a small draft model to guess several tokens ahead, then verifies them all in one parallel pass with the full model, speeding up generation.
Batching (Continuous Batching)
Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish.