Perceptron
The Perceptron (1958) was the first learning system built from an artificial neuron, and the direct ancestor of the neural network training loop.
The Perceptron (1958, Frank Rosenblatt) was the first system built from an artificial neuron that could actually learn its weights from labeled examples, rather than being hand-programmed — a direct, if primitive, ancestor of the gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. training loop used throughout modern deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones.. In 1969, Marvin Minsky and Seymour Papert's book Perceptrons proved a single-layer perceptron can't learn certain simple functions (like XOR), which stalled connectionist research funding for over a decade — until backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network. provided the fix for multi-layer networks.
How it works
A perceptron computes y = step(w · x + b): a weighted sum of the
inputs passed through a hard threshold producing 0 or 1. Learning is an
error-correction rule applied one example at a time — predict, compare
against the label, and if wrong nudge the weights toward the correct
answer with w += lr * (target - prediction) * x. Nothing updates when
the prediction is already right. Rosenblatt's convergence theorem
guarantees this terminates in a finite number of updates if the data
is linearly separable. Note what's absent relative to a modern unit:
the hard threshold has no useful derivative, so there is no
loss functionLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. to differentiate and no way
to push error back through a second layer.
When it breaks
- The decision boundary is a single hyperplane, so any problem that isn't linearly separable — XOR being the canonical case — cannot be learned at all. That's the result Minsky and Papert formalized.
- On non-separable data the update rule never settles; the weights oscillate indefinitely instead of converging to a best-effort fit.
- Stacking perceptrons doesn't help without a differentiable activation functionActivation FunctionAn activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function.. That swap is what makes multi-layer networks trainable and gives them their universal approximationUniversal Approximation TheoremThe universal approximation theorem proves that a feedforward network with even one hidden layer can approximate any continuous function, given enough neurons. property.
- The linear-threshold unit itself never went away: it is exactly one row of a modern layer's matrix multiply.
See also: Neural NetworkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation., BackpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network.
Learn more: History & Landscape · Book: Perceptrons (Minsky & Papert, 1969) · Wikipedia: Perceptron
Mentioned in
Lessons where this comes up in context.
Expert System
An expert system encodes a human expert's domain knowledge as hand-written if-then rules plus an inference engine, the dominant AI approach of the 1970s-80s.
word2vec
word2vec (2013) was a technique for learning dense vector representations of words from raw text, a direct precursor to modern token embeddings.