History

Perceptron

The Perceptron (1958) was the first learning system built from an artificial neuron, and the direct ancestor of the neural network training loop.

The Perceptron (1958, Frank Rosenblatt) was the first system built from an artificial neuron that could actually learn its weights from labeled examples, rather than being hand-programmed — a direct, if primitive, ancestor of the gradient descent training loop used throughout modern deep learning. In 1969, Marvin Minsky and Seymour Papert's book Perceptrons proved a single-layer perceptron can't learn certain simple functions (like XOR), which stalled connectionist research funding for over a decade — until backpropagation provided the fix for multi-layer networks.

How it works

A perceptron computes y = step(w · x + b): a weighted sum of the inputs passed through a hard threshold producing 0 or 1. Learning is an error-correction rule applied one example at a time — predict, compare against the label, and if wrong nudge the weights toward the correct answer with w += lr * (target - prediction) * x. Nothing updates when the prediction is already right. Rosenblatt's convergence theorem guarantees this terminates in a finite number of updates if the data is linearly separable. Note what's absent relative to a modern unit: the hard threshold has no useful derivative, so there is no loss function to differentiate and no way to push error back through a second layer.

When it breaks

  • The decision boundary is a single hyperplane, so any problem that isn't linearly separable — XOR being the canonical case — cannot be learned at all. That's the result Minsky and Papert formalized.
  • On non-separable data the update rule never settles; the weights oscillate indefinitely instead of converging to a best-effort fit.
  • Stacking perceptrons doesn't help without a differentiable activation function. That swap is what makes multi-layer networks trainable and gives them their universal approximation property.
  • The linear-threshold unit itself never went away: it is exactly one row of a modern layer's matrix multiply.

See also: Neural Network, Backpropagation

Learn more: History & Landscape · Book: Perceptrons (Minsky & Papert, 1969) · Wikipedia: Perceptron

Mentioned in

Lessons where this comes up in context.

On this page