Deep Learning
Deep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones.
Deep learning is the branch of machine learningMachine Learning (ML)Machine learning is the practice of writing programs that learn a function from data, via a loss function and an optimizer, instead of following hand-written rules. built on neural networksNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. with many stacked layers ("depth"). Its defining advantage over earlier statistical ML: a deep network learns its own features directly from raw data (pixels, raw text) instead of requiring a person to hand-engineer them first. It became practical around 2012 once large datasets, GPU compute, and training refinements arrived together.
How it works
Depth buys a hierarchy of representations. Early layers of a CNNCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently. respond to edges and colour blobs, middle layers to textures and parts, later layers to whole objects — none of which was specified by hand. The same idea holds for text, where lower layers of a transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. track surface syntax and higher layers carry more abstract structure.
The training loop is unchanged from classical ML: forward pass, compute a lossLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training., backpropagateBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network., step the optimizer. What makes depth work in practice is a set of enabling techniques — residual connectionsResidual ConnectionA residual (skip) connection adds a layer's input back to its output, giving gradients a direct path backward and making very deep networks trainable., normalization layers, careful initializationWeight InitializationWeight initialization sets a neural network's starting parameter values, using schemes like Xavier/He initialization to keep gradients well-scaled from the first training step., and ReLU-family activations — plus GPUsGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires. to make the matrix multiplications fast enough to iterate.
When it breaks
- Data hunger. Deep models need far more labelled examples than classical methods. On a few thousand rows of tabular data, gradient boosted trees usually win outright.
- It fails silently. A bug in preprocessing, a mislabelled class, or a too-high learning rate produces a model that trains without error and simply performs worse — there is no stack trace.
- Shortcut learning. Networks latch onto whatever correlates with the label in your data (watermarks, backgrounds, hospital scanner IDs) and collapse on deployment data that lacks the shortcut.
- Cost and reproducibility. Nondeterministic GPU kernels, seed sensitivity, and long training runs make results hard to reproduce exactly.
See also: Neural NetworkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation., BackpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network., CNNCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently., TransformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.
Learn more: History & Landscape · Wikipedia: Deep learning
Mentioned in
Lessons where this comes up in context.
- Computer VisionConvolutions, pooling, CNNs, transfer learning
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- ML FundamentalsSupervised/unsupervised learning, loss functions, gradient descent
- Neural Networks & BackpropFrom Karpathy's micrograd approach — building a tiny neural net and stepping through forward/backward passes
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
Machine Learning (ML)
Machine learning is the practice of writing programs that learn a function from data, via a loss function and an optimizer, instead of following hand-written rules.
Neural Network
A neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation.