Jürgen Schmidhuber
Co-developed LSTM in 1997, the recurrent architecture that made training on long sequences practical years before transformers existed.
Jürgen Schmidhuber (b. 1963) co-developed, with his student Sepp Hochreiter, the long short-term memory (LSTM) network in 1997 — a variant of the RNNRNN (Recurrent Neural Network)An RNN is a neural network for sequences that processes input one step at a time, carrying a hidden state forward — the main predecessor to transformers. designed specifically to fix the vanishing gradient problemVanishing/Exploding GradientsVanishing/exploding gradients occur when the product of many local gradients across deep layers shrinks toward zero or grows unboundedly, breaking training. that made plain RNNs unable to learn dependencies across long sequences. LSTMs add a gated memory cell that can preserve information across many time steps, letting a network learn, for instance, that a pronoun forty words back determines the correct verb conjugation now.
For roughly two decades — well before transformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. existed — LSTMs were the default architecture for anything sequential: speech recognition, machine translation, and early language modeling all ran on them, including the sequence-to-sequence systems Ilya SutskeverIlya SutskeverCo-authored AlexNet as a student, then sequence-to-sequence learning, then co-founded OpenAI and helped drive the bet that scale would produce GPT-level language models. worked on before transformers displaced recurrence entirely in 2017.
Schmidhuber has also been an outspoken, sometimes contentious voice about priority disputes in the field, frequently arguing that his lab's earlier work anticipated ideas later credited to others, including aspects of attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. mechanisms and generative adversarial networksGAN (Generative Adversarial Network)A GAN trains a generator and a discriminator against each other, the generator learning to fool the discriminator into mistaking its output for real data.. Whatever the scorekeeping, LSTM's practical impact on a generation of NLP and speech systems is not in dispute.
See also: Ilya SutskeverIlya SutskeverCo-authored AlexNet as a student, then sequence-to-sequence learning, then co-founded OpenAI and helped drive the bet that scale would produce GPT-level language models.
Learn more: LLMs · Wikipedia: Jürgen Schmidhuber
Andrej Karpathy
Former Tesla AI director and OpenAI founding member known for building neural networks from scratch on video and coining the phrase "vibe coding."
Demis Hassabis
Co-founded DeepMind and led AlphaGo, AlphaFold, and Gemini — spanning reinforcement learning breakthroughs, a Nobel Prize-winning protein-folding model, and frontier LLMs.