Tooling & The Dev Stack
Languages, frameworks, and where they fit — what you'd actually touch to build and ship a model
The last two lessons covered what a neural networkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. is and how it's trained, in math. This lesson is the translation layer: what does that look like in actual code, on an actual machine, and which tools do you reach for at each stage? If you're a working developer who's used AIAI (Artificial Intelligence)AI is the field of building systems that perform tasks normally requiring human intelligence — reasoning, perception, language, and decision-making. products for years but never built one, this is the map you're missing.
The stack is a tower of abstractions, and each layer exists to hand something specific to the one above it. It's worth seeing the shape of it before walking through the layers one at a time:
Most developers spend their time in the middle two layers, reach down to the driver layer only when something breaks, and reach up to the serving layer once a model actually has users.
The language layer: (almost always) Python
Nearly all model development happens in Python — not because Python is fast (it isn't), but because of what sits underneath it: libraries like NumPy, PyTorchPyTorchPyTorch is the dominant deep learning framework in both research and production, providing tensor computation, GPU dispatch, and automatic differentiation., and TensorFlowTensorFlowTensorFlow is a deep learning framework, dominant in the mid-2010s, still widely used in production and edge deployment (via TensorFlow Lite). expose a Python API while doing the actual number-crunching in compiled C/C++/CUDACUDACUDA is NVIDIA's programming platform that lets frameworks like PyTorch dispatch tensor computation to NVIDIA GPUs. code. You write high-level Python that describes what computation to run; the framework dispatches that computation to fast, low-level kernels. You get Python's ergonomics without paying Python's speed penalty for the parts that matter.
A few other languages show up at specific layers:
- C++/CUDA: the actual matrix-multiplication and attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. kernels that run on the GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires. are written here, not Python. You'd touch this layer directly only if you're writing custom high-performance kernels — most developers never do.
- Rust: an emerging alternative for inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. specifically — projects
like
candle(Hugging FaceHugging FaceHugging Face is an ecosystem — transformers, datasets, and the Hub — providing pretrained models and standardized tooling on top of frameworks like PyTorch.) andburnlet you run models with no Python runtime dependency, useful for edge/embedded deployment. Not used for training at scale. - Julia: popular in some scientific-computing and research circles for its speed-without-C-extensions design, but has nowhere near Python's ecosystem gravity for MLMachine Learning (ML)Machine learning is the practice of writing programs that learn a function from data, via a loss function and an optimizer, instead of following hand-written rules..
Bottom line: learn Python for anything model-development-related. You don't need C++/CUDA unless you're doing performance engineering specifically.
The framework layer: what defines a model in code
A deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones. framework gives you two core things: a way to define a computation graph of tensor operations (a tensor is just a multi-dimensional array of numbers — a single number is a 0-d tensor, a list of numbers is 1-d, a grid like an image is 2-d, and so on; every input, weight, and activation in this course has been a tensor whether or not it was called one) — the forward pass — and automatic differentiation (autograd) — it computes the backward pass (backprop, from the previous lesson) for you, automatically, from the forward pass you wrote. You never hand-derive gradients in practice; the framework does it.
Automatic differentiation is not symbolic calculus and it is not numerical finite differences. It is bookkeeping. As the forward pass runs, the framework builds a tape — a directed graph whose nodes are the primitive operations that were executed, in the order they executed, each storing whatever intermediate values its derivative will need.
Each primitive contributes not its full Jacobian but a vector-Jacobian product: a function that maps an incoming gradient to the outgoing gradient
Replaying the tape backwards and composing these products is exactly the chain rule, and it is why the whole thing is called reverse-mode autodiff. For a loss the cost of one backward pass is a small constant multiple of the forward pass, regardless of how many parameters there are — which is the property that makes training billion-parameter models possible at all.
The price is memory. Every intermediate value the backward pass will need must stay resident until it is consumed, so activation memory scales with depth times batch size. Gradient checkpointing trades that back: discard most intermediates and recompute them during the backward pass, buying a large memory saving for roughly one extra forward pass of compute.
PyTorch builds this tape implicitly as you call operations on tensors that
have requires_grad set. TensorFlow's GradientTape is the same idea with
the name left visible. JAX instead traces your function once and transforms
it, which is why it demands pure functions and no hidden state.
- PyTorch — the dominant framework today, in both research and most production training. Defines computation "eagerly" (operations run immediately, like normal Python), which makes debugging and experimentation fast.
- JAX — Google-associated, built around function transformations
(
grad,jit,vmap) rather than an object-oriented model API. Popular in research and for some of the largest-scale training runs, because its functional style composes unusually well with distributed hardware. Has a steeper learning curve than PyTorch. - TensorFlow/Keras — dominant in the mid-2010s, still widely used in existing production systems and in some mobile/edge deployment paths (via TensorFlow Lite), but has lost most of the research and new-project share to PyTorch over the last several years.
| Framework | Execution style | Best for | Learning curve |
|---|---|---|---|
| PyTorch | Eager — runs immediately, like normal Python | Research and most production training today | Gentlest — the default recommendation |
| JAX | Function transformations (grad, jit, vmap) | Research and some of the largest-scale training runs | Steepest — a different mental model |
| TensorFlow/Keras | Eager by default, graph mode available | Existing production systems, mobile/edge (TF Lite) | Moderate — familiar if you already know it |
The same training step — the "model, loss, gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient." loop from ML Fundamentals — looks different in each, which is a useful way to feel the difference in style:
import torch
pred = model(x) # forward pass
loss = loss_fn(pred, y)
loss.backward() # backprop — autograd computes all gradients
optimizer.step() # gradient descent update
optimizer.zero_grad()Eager and object-oriented: loss.backward() is backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network., done
for you, on a model object that owns its own parameters.
Newer/niche: MLX (Apple, optimized for Apple Silicon, mainly for local experimentation), Candle/Burn (Rust, inference-focused, no Python runtime needed).
Bottom line: default to PyTorch unless you have a specific reason not to. It's what most tutorials, pretrained models, and job postings assume.
The ecosystem layer: you rarely start from scratch
Almost nobody trains a model architecture entirely from raw PyTorch tensors anymore. A layer of libraries sits on top of the framework and handles the repetitive parts:
- Hugging Face
transformers: thousands of pretrained model architectures (the transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. from the previous lessons, and many variants) with a consistent API to load, fine-tune, and run them. For most practical work, this is where you start — not by writing a transformer block from scratch. - Hugging Face
datasetsand the Hugging Face Hub: standardized access to public datasets and pretrained model weights, so you're not hunting down and reformatting data yourself. accelerate/PEFT:acceleratehandles distributing training across multiple GPUs with minimal code changes;PEFTimplements parameter-efficient fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions. methods (like LoRALoRA (Low-Rank Adaptation)LoRA is a parameter-efficient fine-tuning method that trains a small number of additional low-rank parameters instead of updating an entire model's weights. — fine-tuning by training a small number of additional parameters instead of the whole model), which is how most fine-tuning is actually done today given how large models have gotten.- NumPy / pandas: the general-purpose numerical and tabular-data libraries underneath almost everything else in the Python data/ML ecosystem, including parts of PyTorch itself.
Training at scale: distributed training frameworks
A model that fits on one GPU trains the way the code snippet above shows. Training something LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama.-sized requires spreading the model and data across many GPUs (sometimes thousands), which introduces its own tooling layer:
- PyTorch FSDP (Fully Sharded Data Parallel) and DeepSpeed (Microsoft): split a model's parameters, gradients, and optimizer state across multiple GPUs, so a model too large to fit on one GPU's memory can still be trained.
- Megatron-LM (NVIDIA): specifically built for training very large transformer models efficiently across many GPUs, combining several parallelism strategies at once.
- Ray: general-purpose distributed-computing framework, commonly used to orchestrate large training and data-processing jobs across a cluster, not ML-specific.
You won't touch these directly for small projects — they matter once you're training something too large for a single machine, which for most developers experimenting or fine-tuning is not the common case.
The reason this tooling exists at all is that model weights are rarely what fills a GPU. Training also has to hold gradients, optimizer state, and the stored activations autograd needs — and those together usually dwarf the weights themselves.
Take a 7-billion-parameter model, fine-tuned in the standard mixed-precision setup with Adam. Count 2 bytes per fp16 number and 4 per fp32:
- Weights, fp16: 7B × 2 = 14 GB
- Gradients, fp16: 7B × 2 = 14 GB
- Adam state, fp32 master weights + two moments: 7B × 4 × 3 = 84 GB
That's ~112 GB before a single activation is stored — on a GPU with 80 GB of memory. The optimizer state alone is six times the size of the fp16 weights.
This is why the two most common fixes are (a) shard those tensors across GPUs with FSDP or DeepSpeed, or (b) don't create most of them at all: LoRA freezes the base weights, so gradients and Adam state exist only for the small adapter, collapsing the last two lines to near zero and leaving the 14 GB of frozen weights as the dominant term.
For parameters, a training step's memory splits into four terms:
The first three are fixed by the model and optimizer and do not depend on batch size. In mixed-precision training with Adam, and — four bytes each for the fp32 master copy of the weights, the first moment , and the second moment . Plain SGD without momentum has , which is a large part of why it is still used when memory is the binding constraint.
The activation term behaves completely differently. It scales roughly as
for batch size , sequence length , and hidden size — so it is the only term you control at runtime. Attention adds a term that grows with unless a fused implementation avoids materialising the full score matrix, which is what FlashAttention-style kernels are for.
The practical consequence: an out-of-memory error at step 1 is a weights-plus-optimizer problem and needs sharding, quantizationQuantizationQuantization reduces a model's numerical precision (e.g. 16-bit to 4-bit) to shrink memory footprint and speed up inference, at some cost to accuracy., or a different optimizer. An out-of-memory error part-way into a run is almost always an activation problem and is fixed by a smaller batch, a shorter sequence, or gradient checkpointing (discarding activations and recomputing them on the backward pass instead of storing them, as covered above).
Keeping track of what you're doing: experiment tooling
Training runs take hours to weeks, and you'll run many variations (different learning rates, architectures, data). Two categories of tooling exist specifically to keep that manageable:
- Experiment tracking: Weights & Biases and MLflow log metrics, hyperparameters (the settings you chose before the run — learning rate, batch size, and so on), and outputs across runs so you can compare them later instead of relying on memory or scattered notebooks.
- Notebooks vs. scripts: Jupyter notebooks are common for exploration and visualization; actual training jobs (especially anything long-running or distributed) are almost always run as plain Python scripts, since notebooks don't handle multi-hour, multi-GPU, restart-safe jobs well.
The hardware layer
None of the above runs without hardware underneath it:
- NVIDIA GPUs + CUDA: the dominant combination by a wide margin. CUDA is NVIDIA's programming platform that lets frameworks like PyTorch dispatch computation to the GPU; cuDNN provides optimized implementations of common deep learning operations on top of it.
- TPUs (Tensor Processing Units, Google): custom hardware built specifically for the matrix operations deep learning needs; used heavily inside Google and paired closely with JAX.
- AMD (ROCm): NVIDIA's real alternative on the hardware side, with a growing but still much smaller software ecosystem in comparison.
Modern accelerators are also much faster at low-precision arithmetic than at the 32-bit floats you'd write by default, which is why virtually all training today is mixed precision — most operations in a 16-bit format, a few numerically delicate ones kept in fp32. The two 16-bit formats differ in how they spend their bits:
| Format | Exponent bits | Mantissa bits | Practical consequence |
|---|---|---|---|
| fp32 | 8 | 23 | Baseline range and precision |
| fp16 | 5 | 10 | Narrow range; needs loss scaling |
| bf16 | 8 | 7 | fp32 range, coarser steps; no scaling needed |
A floating-point number is , so the exponent bits buy range and the mantissa bits buy precision. fp16 keeps only 5 exponent bits, which puts its smallest normal positive value around , with subnormals reaching to roughly before hitting zero.
That is a real problem in the backward pass. Gradients in deep networks are routinely smaller than , and anything below the subnormal floor doesn't become small — it becomes exactly , and that zero then propagates through the chain rule and silently kills the update for everything upstream of it.
Loss scaling is the fix, and it works because differentiation is linear. Multiply the loss by a constant before the backward pass and every gradient is multiplied by the same :
So small gradients get lifted into fp16's representable range, and the optimizer divides by again before applying the update. Nothing about the mathematics changes — only which numbers survive the trip. Frameworks pick dynamically: raise it while steps are clean, and on any step where a gradient overflows to infinity or NaN, throw that step away and halve .
bf16 sidesteps the whole mechanism by keeping fp32's 8 exponent bits and spending the savings on precision instead — same dynamic range as fp32, so no underflow and no scaling required, at the cost of a coarser mantissa. This is why bf16 is the default on hardware that supports it and fp16 with loss scaling is what you use on hardware that doesn't. In either case the optimizer still keeps an fp32 master copy of the weights, because a 16-bit weight plus a much smaller 16-bit update frequently rounds to the unchanged weight.
After training: the deployment/inference layer
Once you have trained weights, running them efficiently in production uses a different set of tools than training did (see Inference & Serving for the concepts these implement):
- vLLM: a widely used open-source LLM serving engine, built around efficient KV-cache management and continuous batchingBatching (Continuous Batching)Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish..
- TensorRT-LLM (NVIDIA) and Triton Inference Server: NVIDIA's optimized inference stack, commonly used in enterprise deployments.
- ONNX Runtime: runs models exported to the ONNX format (a framework-independent representation), useful for deploying a model outside the framework it was trained in.
- llama.cpp / GGUF, Ollama: the path for running models locally on a laptop or edge device with no GPU cluster required — quantization (previous lesson) is central to making this practical.
A concrete starting stack
If you're a developer who wants to actually build something rather than just survey the landscape, a reasonable default stack today:
Language: Python
Framework: PyTorch
Pretrained models/data: Hugging Face (transformers, datasets, Hub)
Fine-tuning: PEFT (LoRA) + accelerate, on a single or few GPUs
Experiment tracking: Weights & Biases
Local inference: Ollama or llama.cpp, for trying things without a GPU cluster
Production inference: vLLM, once you're serving real trafficThis isn't the only valid stack — JAX is a legitimate alternative at the framework layer, especially for research; TensorFlow still runs plenty of production systems — but it's the one with the largest community, the most tutorials, and the most pretrained models available out of the box, which makes it the highest-leverage starting point.
Recap
Python is the language you write; PyTorch (usually) is the framework that gives you autograd and GPU dispatch; Hugging Face is the ecosystem layer that means you rarely start from a blank file; distributed-training and experiment-tracking tools exist for when a project outgrows a single GPU or a single run; and a separate deployment stack (vLLM and friends) takes over once you're serving trained weights instead of producing them. The next lesson returns to architecture — how CNNsCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently. apply everything from Neural Networks & Backprop to images specifically.