GPU (Graphics Processing Unit)
GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires.
A GPU (Graphics Processing Unit), originally built for rendering graphics, turned out to be extremely well-suited to the large parallel matrix multiplications deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones. needs — one of the three factors (alongside data and algorithms) that made deep learning practical starting around 2012. NVIDIA GPUs paired with CUDACUDACUDA is NVIDIA's programming platform that lets frameworks like PyTorch dispatch tensor computation to NVIDIA GPUs. are the dominant combination for both training and inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering.; TPUs (Google) and AMD's ROCm are the main alternatives.
How it works
A CPU devotes most of its die to control logic and cache to make a few sequential threads fast. A GPU inverts that: thousands of simple cores run the same instruction across many data elements at once, which maps directly onto the dense matrix multiplies inside a neural networkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation.. Modern accelerators also carry dedicated matrix units (NVIDIA calls them Tensor Cores) that multiply small tiles in reduced precision and accumulate in higher precision.
Two numbers govern performance: arithmetic throughput and memory bandwidth. Training large models is usually compute-bound, while autoregressive decoding in an LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. is memory-bound — every generated token re-reads the weights, so batchingBatching (Continuous Batching)Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish. requests is what recovers utilization.
When it breaks
- Memory capacity is the hard ceiling. Training needs room for weights, activations, gradients, and optimizer state — typically several times the parameter count in bytes. This is what forces sharding long before compute becomes the limit.
- Small or irregular workloads waste the device. Tiny batches, many small kernel launches, or heavy Python-side logic leave the GPU idle while it waits on the host.
- Bandwidth, not FLOPs, is often the bottleneck. A kernel that reads more bytes than it does arithmetic will hit nowhere near peak throughput no matter how the code is tuned.
- Portability is not free. Code tuned for one vendor's stack rarely transfers cleanly; custom kernels and quantizationQuantizationQuantization reduces a model's numerical precision (e.g. 16-bit to 4-bit) to shrink memory footprint and speed up inference, at some cost to accuracy. formats are commonly tied to a specific hardware generation.
See also: CUDACUDACUDA is NVIDIA's programming platform that lets frameworks like PyTorch dispatch tensor computation to NVIDIA GPUs., InferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering.
Learn more: History & Landscape · Tooling & The Dev Stack
Mentioned in
Lessons where this comes up in context.
- Attention & TransformersThe core architecture behind modern AI
- Computer VisionConvolutions, pooling, CNNs, transfer learning
- Generative ModelsGANs, VAEs, and diffusion models — how AI generates new images, audio, and video, as opposed to classifying or understanding existing content
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- Inference & ServingBatching, KV caching, quantization, latency/throughput tradeoffs
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- Optimization & Training DynamicsWhy plain gradient descent isn't what actually trains modern models — Adam's mechanics, why warmup exists, how learning-rate schedules are shaped, and what batch size actually trades off at scale
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model