Tooling & Hardware

CUDA

CUDA is NVIDIA's programming platform that lets frameworks like PyTorch dispatch tensor computation to NVIDIA GPUs.

CUDA is NVIDIA's programming platform that lets frameworks like PyTorch dispatch computation to NVIDIA GPUs. Deep learning frameworks expose a high-level Python API while their actual matrix-multiplication and attention kernels run as compiled CUDA code underneath. cuDNN provides optimized implementations of common deep learning operations on top of CUDA. NVIDIA GPUs plus CUDA form the dominant hardware/software combination for both training and serving models.

How it works

A CUDA program splits work into kernels — functions that run once per thread, with threads grouped into blocks and blocks into a grid. A kernel launch queues work onto a stream, an ordered queue the GPU executes asynchronously while the CPU keeps running Python. That asynchrony is why timing GPU code requires an explicit torch.cuda.synchronize().

  • Memory is explicit: tensors live in device memory, and host/device copies cross the PCIe bus, which is far slower than on-device bandwidth.
  • Libraries do the heavy lifting: cuBLAS for matrix multiply, cuDNN for convolutions and normalization, NCCL for multi-GPU collectives used in distributed training.
  • Tensor Cores on recent architectures execute mixed-precision matrix multiplies far faster than standard FP32 paths.

When it breaks

  • Version mismatch. The driver, the CUDA runtime, cuDNN, and the framework build must be mutually compatible. A framework built for one CUDA major version will refuse to initialize — or fail at the first kernel launch — against an older driver.
  • Silent host/device transfers. Calling .item(), .numpy(), or printing a tensor forces a synchronize and a copy back to the CPU. Inside a training loop this serializes the pipeline and can dominate step time.
  • Fragmented memory. CUDA out of memory often fires with free memory still reported, because the caching allocator cannot find a contiguous block. Varying sequence lengths during inference make this worse.
  • Async errors surface late. Because launches are queued, a fault is frequently reported at an unrelated later line unless you set CUDA_LAUNCH_BLOCKING=1.

See also: GPU, PyTorch

Learn more: Tooling & The Dev Stack · NVIDIA CUDA docs

Mentioned in

Lessons where this comes up in context.

On this page