CUDA
CUDA is NVIDIA's programming platform that lets frameworks like PyTorch dispatch tensor computation to NVIDIA GPUs.
CUDA is NVIDIA's programming platform that lets frameworks like PyTorchPyTorchPyTorch is the dominant deep learning framework in both research and production, providing tensor computation, GPU dispatch, and automatic differentiation. dispatch computation to NVIDIA GPUsGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires.. Deep learning frameworks expose a high-level Python API while their actual matrix-multiplication and attention kernels run as compiled CUDA code underneath. cuDNN provides optimized implementations of common deep learning operations on top of CUDA. NVIDIA GPUs plus CUDA form the dominant hardware/software combination for both training and serving models.
How it works
A CUDA program splits work into kernels — functions that run once per
thread, with threads grouped into blocks and blocks into a grid. A
kernel launch queues work onto a stream, an ordered queue the GPU
executes asynchronously while the CPU keeps running Python. That
asynchrony is why timing GPU code requires an explicit
torch.cuda.synchronize().
- Memory is explicit: tensors live in device memory, and host/device copies cross the PCIe bus, which is far slower than on-device bandwidth.
- Libraries do the heavy lifting:
cuBLASfor matrix multiply,cuDNNfor convolutions and normalization,NCCLfor multi-GPU collectives used in distributed trainingDistributed Training (FSDP, DeepSpeed, Megatron-LM)Distributed training frameworks split a model's parameters, gradients, and optimizer state across many GPUs so models too large for one GPU can still be trained.. - Tensor Cores on recent architectures execute mixed-precision matrix multiplies far faster than standard FP32 paths.
When it breaks
- Version mismatch. The driver, the CUDA runtime, cuDNN, and the framework build must be mutually compatible. A framework built for one CUDA major version will refuse to initialize — or fail at the first kernel launch — against an older driver.
- Silent host/device transfers. Calling
.item(),.numpy(), or printing a tensor forces a synchronize and a copy back to the CPU. Inside a training loop this serializes the pipeline and can dominate step time. - Fragmented memory.
CUDA out of memoryoften fires with free memory still reported, because the caching allocator cannot find a contiguous block. Varying sequence lengths during inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. make this worse. - Async errors surface late. Because launches are queued, a fault
is frequently reported at an unrelated later line unless you set
CUDA_LAUNCH_BLOCKING=1.
See also: GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires., PyTorchPyTorchPyTorch is the dominant deep learning framework in both research and production, providing tensor computation, GPU dispatch, and automatic differentiation.
Learn more: Tooling & The Dev Stack · NVIDIA CUDA docs
Mentioned in
Lessons where this comes up in context.
Distributed Training (FSDP, DeepSpeed, Megatron-LM)
Distributed training frameworks split a model's parameters, gradients, and optimizer state across many GPUs so models too large for one GPU can still be trained.
GPU (Graphics Processing Unit)
GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires.