Universal Approximation Theorem
The universal approximation theorem proves that a feedforward network with even one hidden layer can approximate any continuous function, given enough neurons.
The universal approximation theorem proves that a feedforward neural networkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. with even a single hidden layer can approximate any continuous function to arbitrary precision, given enough neurons in that layer. It explains why stacking simple weighted-sum-plus-nonlinearity units can represent complex functions at all — though it says nothing about whether such a network is practically learnable via gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient., which is why depth (not just width) matters in practice.
How it works
The classic result (Cybenko 1989, Hornik 1991) is an existence proof.
For any continuous function f on a closed, bounded input region and any
tolerance ε > 0, there exist weights for a one-hidden-layer network
such that the network stays within ε of f everywhere on that region.
The intuition: each hidden unit contributes a shifted, scaled copy of the activation functionActivation FunctionAn activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function., and a weighted sum of enough such bumps can trace any continuous curve — much like a piecewise-constant approximation, but smooth. Hornik's version shows the property comes from having any non-polynomial activation, not from sigmoid specifically.
Later work proves a dual form: networks of bounded width but unbounded depth are also universal.
When it breaks
- It says nothing about how many neurons. The required width can grow exponentially with input dimension, so "one layer suffices" is not an architectural recommendation.
- Existence ≠ trainability. The theorem guarantees a weight setting exists; it gives no reason to believe backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network. will find it from a random start.
- No generalization claim. Approximation is on the region you approximate over. Outside it, behaviour is unconstrained, and approximating the training points is exactly overfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data..
- Misused as an argument against depth. Deep networks win because they express compositional structure with far fewer parameters, which the theorem does not address.
See also: Neural NetworkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation., Activation FunctionActivation FunctionAn activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function.
Learn more: Neural Networks & Backprop · Wikipedia: Universal approximation theorem
Mentioned in
Lessons where this comes up in context.
Residual Connection
A residual (skip) connection adds a layer's input back to its output, giving gradients a direct path backward and making very deep networks trainable.
Loss Function
A loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training.