PyTorch
PyTorch is the dominant deep learning framework in both research and production, providing tensor computation, GPU dispatch, and automatic differentiation.
PyTorch is the dominant deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones. framework today, in both research and most production model training. It defines computation "eagerly" (operations run immediately, like normal Python), which makes debugging and experimentation fast, and provides autograd — automatic computation of gradients via backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network. — so training code rarely requires hand-deriving gradients. It's the default starting point recommended for most model development work today.
How it works
Every tensor carries a device and a dtype, and operations run
immediately on that device. When a tensor has requires_grad=True,
PyTorch records each operation into a dynamic graph as it executes;
calling .backward() on a scalar lossLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. walks
that graph in reverse, accumulating gradients into each leaf tensor's
.grad. The optimizer then applies them, and optimizer.zero_grad()
clears the accumulation for the next step.
Models subclass nn.Module, which tracks parameters and submodules so
.to("cuda") moves the whole tree onto a GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires. at once.
For speed, torch.compile traces the module and hands it to a backend
that fuses kernels ahead of time, and torch.no_grad() skips graph
construction entirely during inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering..
When it breaks
- Gradients accumulate by default. Forgetting
zero_grad()sums gradients across steps. Training still runs; it just converges to something wrong. - Device mismatches. A model on the GPU and a batch left on the CPU
raises at the first op. The reverse case — a stray
.cpu()inside the loop — raises nothing and quietly halves throughput. - Holding on to the graph. Appending a live loss tensor to a list
keeps its whole autograd graph alive, so memory grows each step until
OOM. Use
.detach()or.item()for logging. - Graph breaks and recompiles.
torch.compilefalls back to eager on unsupported Python, and changing input shapes forces a recompile, so the speedup can silently disappear.
See also: TensorFlowTensorFlowTensorFlow is a deep learning framework, dominant in the mid-2010s, still widely used in production and edge deployment (via TensorFlow Lite)., JAXJAXJAX is a Google-associated deep learning framework built around function transformations (grad, jit, vmap), popular in research and large-scale training., Hugging FaceHugging FaceHugging Face is an ecosystem — transformers, datasets, and the Hub — providing pretrained models and standardized tooling on top of frameworks like PyTorch.
Learn more: Tooling & The Dev Stack · PyTorch official docs
Mentioned in
Lessons where this comes up in context.
Red Teaming
Red teaming is the practice of deliberately probing a deployed model for harmful, unsafe, or exploitable behavior before and after release.
TensorFlow
TensorFlow is a deep learning framework, dominant in the mid-2010s, still widely used in production and edge deployment (via TensorFlow Lite).