Neural Network
A neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation.
A neural network is a model built from layers of neurons — units that compute a weighted sum of their inputs, add a bias, and pass the result through a nonlinear activation function (e.g. ReLU). Stacking layers lets the network represent increasingly complex functions; its parameters (weights and biases) are trained via gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient., with backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network. computing the gradients efficiently. A basic fully-connected network is called an MLP (multi-layer perceptron); CNNsCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently. and transformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. are neural networks with additional structure suited to images and sequences respectively.
How it works
A fully-connected layer is one matrix multiply plus a bias plus a
nonlinearity: h = f(x @ W + b), with W of shape
(in_features, out_features). An MLP is just several of these composed,
so the whole network is a single differentiable function from input to
output.
Training runs in three phases per step:
- Forward: push a batch through the layers to get predictions and a lossLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training..
- Backward: backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network. walks the
graph in reverse to get
dL/dWfor every weight. - Update: the optimizer applies
W := W - lr * dL/dW.
Width (neurons per layer) and depth (number of layers) both add capacity, but depth is what lets features compose — see the universal approximation theoremUniversal Approximation TheoremThe universal approximation theorem proves that a feedforward network with even one hidden layer can approximate any continuous function, given enough neurons. for why width alone is theoretically sufficient yet practically insufficient.
When it breaks
- Shape and dtype errors dominate early debugging. A transposed
matrix or a silently broadcast
(N, 1)against(N,)produces a loss that trains but means nothing. - Unscaled inputs. Features spanning wildly different magnitudes make the loss surface badly conditioned; standardize inputs before training.
- Depth without support structure. Past roughly a dozen plain layers training stalls from vanishing gradientsVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training. unless you add residual connectionsResidual ConnectionA residual (skip) connection adds a layer's input back to its output, giving gradients a direct path backward and making very deep networks trainable. and normalization.
- Capacity mistaken for skill. A network large enough to memorize the training set will do exactly that; watch validation loss, not training loss.
See also: BackpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network., Gradient DescentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient., Deep LearningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones.
Learn more: Neural Networks & Backprop · Wikipedia: Neural network
Mentioned in
Lessons where this comes up in context.
- Attention & TransformersThe core architecture behind modern AI
- Computer VisionConvolutions, pooling, CNNs, transfer learning
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- ML FundamentalsSupervised/unsupervised learning, loss functions, gradient descent
- Neural Networks & BackpropFrom Karpathy's micrograd approach — building a tiny neural net and stepping through forward/backward passes
- Reinforcement LearningMDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
Deep Learning
Deep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones.
Activation Function
An activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function.