History & Landscape
Symbolic AI to expert systems to statistical ML to deep learning to the LLM era
"AIAI (Artificial Intelligence)AI is the field of building systems that perform tasks normally requiring human intelligence — reasoning, perception, language, and decision-making." has meant different things at different times. Understanding the sequence of ideas — and why each one hit a wall — makes the current moment (large language modelsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama.) look less like magic and more like the next step in a long argument about what "intelligence" even is, computationally.
Five eras, and what pushed the field out of each one:
1950s–1960s
Symbolic AI
Hand-written rules and logic manipulating hand-written representations of the world. Worked well for narrow, well-defined problems — but the rules didn't generalize beyond them.
1970s–1980s
Expert systems
Scaled the rules up: a human expert's domain knowledge encoded as a large
if-thenknowledge base. Hit a knowledge bottleneck — every new domain needed a new expert and a new rule base — and triggered the AI winters.1990s–2000s
Statistical ML
A mindset shift: a model is a function fit to data, not a set of rules written by a person. Still needed hand-engineered features, which capped accuracy on messy, high-dimensional data.
2010s
Deep learning
Networks learn their own features directly from raw data, removing the feature-engineering bottleneck — once data, GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires. compute, and training refinements arrived together.
2017–present
Transformers & LLMs
AttentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer.-based architectures that scale unusually predictably: bigger models, trained on more text, keep getting better in a measurable, extrapolatable way.
Before "AI" had a name: 1940s–1950s origins
Two threads that the rest of this history keeps returning to both start before the field had a name. In 1943, Warren McCullochWarren McCullochNeurophysiologist who, with Walter Pitts, proposed the first mathematical model of a neuron in 1943 — the ancestor of every neural network built since. and Walter PittsWalter PittsSelf-taught logician who, with Warren McCulloch, co-authored the 1943 paper modeling neurons as logic gates — the mathematical starting point for every neural network. described a simplified mathematical model of a neuron — binary inputs, weighted, thresholded — that could compute logical functions. It's the direct conceptual ancestor of the neurons in Neural Networks & Backprop, though nobody yet had a way to train one, only to hand-construct it.
In 1950, Alan TuringAlan TuringMathematician who formalized what "computation" means and, in 1950, asked whether a machine could think — the question the entire field still answers to. published "Computing Machinery and Intelligence," proposing what's now called the Turing Test: rather than define "thinking" philosophically, judge a machine by whether its conversation is indistinguishable from a human's. It's a pragmatic sidestep that still shapes how the field talks about capability today.
The field got its name and founding moment in 1956, at the Dartmouth Summer Research Project, organized by John McCarthyJohn McCarthyCoined the term "artificial intelligence" in 1955, organized the field's founding 1956 Dartmouth workshop, and invented Lisp. (who coined the term "artificial intelligence"), Marvin MinskyMarvin MinskyCo-founder of the MIT AI Lab and a founding figure of the field — his 1969 book with Seymour Papert exposed the perceptron's limits and helped trigger the first AI winter., Nathaniel Rochester, and Claude Shannon. The proposal's premise — that "every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it" — is the optimistic bet the entire field has been testing ever since.
1950s–1960s: Symbolic AI and the first connectionist wave
Two competing bets on how to build that machine emerged almost immediately, and the tension between them recurs throughout this history.
Symbolic AI bet that intelligence is symbol manipulation: if you can represent knowledge as symbols and rules, and manipulate those symbols with logic, you get reasoning. Early programs proved geometry theorems and played checkers this way — hand-written rules operating on hand-written representations of the world. This approach is often called GOFAI (Good Old-Fashioned AI), and it worked surprisingly well for narrow, well-defined problems (chess endgames, symbolic algebra).
Connectionism bet the opposite: that intelligence emerges from simple units (like McCulloch-Pitts neurons) connected together and adjusted by experience, rather than programmed with explicit rules. Frank RosenblattFrank RosenblattPsychologist who built the perceptron in 1958, the first neural network that learned its own weights from data rather than having them set by hand.'s PerceptronPerceptronThe Perceptron (1958) was the first learning system built from an artificial neuron, and the direct ancestor of the neural network training loop. (1958) was the first such system that could actually learn its weights from labeled examples — a direct, if primitive, ancestor of the gradient-descent training loop in ML Fundamentals.
The Perceptron's momentum stalled hard in 1969, when Marvin Minsky and Seymour Papert published Perceptrons, a rigorous analysis proving that a single-layer perceptron cannot learn certain simple functions (the canonical example is XOR — output true if exactly one of two inputs is true). The book didn't claim multi-layer networks were hopeless, but it was widely read as a case against neural networksNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. generally, and connectionist research funding collapsed for over a decade — the field's center of gravity swung firmly back to symbolic AI.
A single-layer perceptron computes a thresholded weighted sum of its inputs:
The set of inputs it maps to 1 is therefore always one side of a hyperplane — in two dimensions, one side of a straight line. A perceptron can only represent linearly separable functions.
XOR is not one. With inputs and weights , the four required outputs give four inequalities:
Adding the middle two gives , and since this forces — contradicting the last inequality. No choice of weights works.
The fix is a second layer: one hidden unit detecting and another detecting make XOR linearly separable in the hidden layer's coordinates. The representational gap was never the problem — training the extra layer was.
The irony worth sitting with: the fix (stacking multiple layers, trained via backpropagationBackpropagationBackpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, by applying the chain rule backward through the network.) existed in principle within another couple of decades, and this same connectionist thread — not symbolic AI — is what underlies every architecture covered later in this course.
1970s–1980s: Expert Systems
The response to GOFAI's limits wasn't to abandon rules, but to scale them
up: expert systemsExpert SystemAn expert system encodes a human expert's domain knowledge as hand-written if-then rules plus an inference engine, the dominant AI approach of the 1970s-80s. encoded a human expert's domain knowledge as a large
base of if-then rules, plus an inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. engine to chain them together.
Systems like MYCIN (medical diagnosis) and XCON (configuring computer
orders) were commercially deployed and genuinely useful.
They also revealed the core problem with hand-coded knowledge: it doesn't generalize and it doesn't scale. Every new domain needed a new expert to interview and a new rule base to hand-write and maintain. This bottleneck — plus overpromising relative to what the systems could deliver — led to funding pullbacks now called the AI winters (mid-1970s, late 1980s–90s; the Perceptrons fallout above was effectively the first one, for connectionism specifically).
Connectionism itself got a working fix during this same period, even though it took years to gain traction: a 1986 paper by David RumelhartDavid RumelhartCognitive scientist who, with Geoffrey Hinton and Ronald Williams, popularized backpropagation in 1986 — the algorithm that made multi-layer neural networks trainable., Geoffrey HintonGeoffrey HintonKnown as a godfather of deep learning — co-authored backpropagation in 1986, co-invented the Boltzmann machine and dropout, and co-authored AlexNet, the 2012 result that restarted the field., and Ronald Williams popularized backpropagation (the algorithm covered in depth in Neural Networks & Backprop) as a practical way to train multi-layer networks — directly answering the single-layer limitation Minsky and Papert had proven. It would still take another two and a half decades, and the arrival of large datasets and GPU compute, for multi-layer networks to decisively outperform other approaches in practice.
The 1986 Rumelhart–Hinton–Williams paper popularized backpropagation; it did not invent it. The underlying method — reverse-mode automatic differentiation — was published by Seppo Linnainmaa in 1970, and Paul Werbos proposed applying it to neural networks in his 1974 PhD thesis.
What the method exploits is an asymmetry in the chain rule. For a network with parameters and a scalar loss, computing every partial derivative by perturbing one parameter at a time costs forward passes. Reverse-mode instead sweeps backwards once, reusing shared subexpressions, and gets all gradients for a cost proportional to a single forward pass:
For a model with parameters, that is the difference between trainable and impossible. The repeated rediscoveries are a recurring pattern in this history: the idea existed years before the field had the data and compute to make it pay off.
1990s–2000s: Statistical Machine Learning
The field's center of gravity shifted from "encode what a human expert knows" to "learn patterns from data." Instead of writing rules, you write an algorithm that finds structure in examples: spam filters that learn word frequencies from labeled email, systems that learn to rank search results from click data.
This era produced ideas still in daily use — decision trees, support vector machines, random forests, logistic regression — and, critically, a mindset shift: a model is a function fit to data, not a set of rules written by a person. That mindset is the direct ancestor of everything that follows in this course.
The catch: these methods mostly needed hand-engineered features. To classify images, someone had to first decide what numbers to extract from pixels (edges, corners, color histograms) before a statistical model could use them. Feature engineering was itself a specialized skill, and it capped how well these systems could do on messy, high-dimensional data like raw pixels or raw text.
2010s: Deep Learning
Deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones. removed the feature-engineering bottleneck: instead of a person deciding what features matter, a neural network learns its own features directly from raw data, layer by layer. This wasn't a new mathematical idea — the core building blocks trace back to the 1980s and earlier (see Neural Networks & Backprop) — it became practical once three things arrived at once:
- Data: internet-scale labeled datasets (ImageNet, for images)
- Compute: GPUs, originally built for rendering graphics, turned out to be extremely good at the matrix multiplications neural networks need
- Algorithms: refinements to training (better initialization, activation functionsActivation FunctionAn activation function is the nonlinearity applied after a neuron's weighted sum, without which stacked layers would collapse into one linear function., regularizationRegularizationRegularization is any technique that trades some training-data fit for better generalization, fighting overfitting on purpose.) that made deep networks actually trainable instead of stalling out
The data piece has its own named milestone: ImageNet, a dataset of over a million labeled images across a thousand categories, assembled by Fei-Fei LiFei-Fei LiBuilt ImageNet, the large labeled dataset that made the 2012 deep learning breakthrough in computer vision possible in the first place.'s group starting in 2009 specifically because existing datasets were too small to reveal what large-scale learning could do. The annual ImageNet competition became deep learning's proving ground.
The turning point most historians point to is 2012: AlexNet (Alex KrizhevskyAlex KrizhevskyBuilt AlexNet in 2012 with Ilya Sutskever and Geoffrey Hinton, the deep convolutional network whose ImageNet win is widely credited with kicking off the deep learning boom., Ilya SutskeverIlya SutskeverCo-authored AlexNet as a student, then sequence-to-sequence learning, then co-founded OpenAI and helped drive the bet that scale would produce GPT-level language models., and Geoffrey Hinton), a deep convolutional network trained on GPUs, beat the previous best image-classification approaches on ImageNet by a wide margin — roughly halving the error rate of the next-best entry in one year. Over the next few years, deep learning took over computer vision, then speech recognition, then — via a 2013 technique called word2vecword2vecword2vec (2013) was a technique for learning dense vector representations of words from raw text, a direct precursor to modern token embeddings. that learned dense vector representations of words from raw text (a direct precursor to the embeddingsEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space. covered in LLMs) — increasingly language, setting the stage for transformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs..
2017–present: Attention, Transformers, and the LLM Era
A 2017 paper, "Attention Is All You Need," introduced the transformer, an architecture built entirely around a mechanism called self-attention (covered in depth in Attention & Transformers). Transformers turned out to scale unusually well: bigger models, trained on more text, kept getting better in fairly predictable ways — an observation formalized as scaling lawsScaling LawsScaling laws are empirical relationships between a model's loss and its parameter count, dataset size, and compute budget, used to plan how large to train a new model..
Kaplan et al. (2020) found that transformer language-model loss falls as a power law in each of model size , dataset size , and training compute , provided the other two are not the bottleneck:
with small exponents — roughly for parameters. A power law plots as a straight line on log–log axes, which is what makes it useful: measure a few small runs, extrapolate the line, and predict the loss of a run you have not yet paid for.
The small exponent is also the catch. Because falls with , halving the loss gap requires increasing by a factor of . Progress is predictable and brutally expensive at the same time — which is exactly the bet the LLM era is making.
Hoffmann et al. (2022) later revised the allocation advice: for a fixed compute budget, earlier models were far too large for the amount of data they were trained on, and parameters and tokens should be scaled roughly in proportion.
That predictability is what justified training ever-larger language models, culminating in the systems now called LLMs (large language models) — GPT, Claude, Gemini, Llama, and others. The recipe that distinguishes this era from everything before it:
- Pretrain a large transformer on a huge, mostly unlabeled text corpus to predict the next token
- Fine-tune it to follow instructions and match human preferences (often via RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward. — reinforcement learning from human feedback)
- Deploy it as a general-purpose system usable for translation, coding, summarization, reasoning, and more — without task-specific training for each one
That last point is the real break from everything earlier in this history: prior eras built one system per task. The LLM era builds one system and prompts it into many tasks.
The public inflection point was ChatGPT's release in November 2022 — not a new architecture or a research breakthrough over what preceded it (GPT-3 had existed since 2020), but the moment this capability became a free-form conversational product anyone could use, which is what made "LLM" a household concept rather than a research term. The GPT-1 → GPT-2 → GPT-3 progression that led up to it is itself a case study in the scaling-laws bet: each generation was substantially larger and trained on more data than the last, with capability increasing in tandem more reliably than most researchers expected going in.
The Mark I Perceptron (1958) adjusted its weights with physical potentiometers — on the order of 10² tunable numbers. GPT-3 (2020) has 1.75 × 10¹¹ parameters.
Roughly a billion-fold increase across 62 years. Spread evenly, that is a doubling roughly every two years — which is to say the field's headline number tracked Moore's law for six decades, and the visible capability jumps came when that curve crossed thresholds, not when it changed slope.
The dataset side moved similarly: the Perceptron learned from hundreds of hand-fed images, ImageNet from ~10⁶, and GPT-3 from ~10¹¹ tokens.
Where this leaves us
Each era didn't so much get "replaced" as get subsumed: statistical MLMachine Learning (ML)Machine learning is the practice of writing programs that learn a function from data, via a loss function and an optimizer, instead of following hand-written rules. still runs most fraud-detection and recommendation systems in production; expert-system-style rule engines still gate business logic; deep learning underlies vision and speech; transformers now dominate language and increasingly other modalities too. The rest of this course follows that stack bottom-up — starting with the statistical-learning foundations everything else builds on.
The through-line is a steady retreat of hand-written work: each era moves one more thing from specified by a person to learned from data.
| Era | Human supplies | System learns |
|---|---|---|
| Symbolic AI | Rules and representations | Nothing |
| Expert systems | Domain rule base | Which rules to chain |
| Statistical ML | Features and model class | Parameters |
| Deep learning | Architecture and labels | Features and parameters |
| LLMs | Architecture and objective | Features, parameters, tasks |