Ilya Sutskever
Co-authored AlexNet as a student, then sequence-to-sequence learning, then co-founded OpenAI and helped drive the bet that scale would produce GPT-level language models.
Ilya Sutskever was a graduate student under Geoffrey HintonGeoffrey HintonKnown as a godfather of deep learning — co-authored backpropagation in 1986, co-invented the Boltzmann machine and dropout, and co-authored AlexNet, the 2012 result that restarted the field. when he co-authored AlexNetAlexNetAlexNet is the 2012 deep convolutional network that won the ImageNet competition by a wide margin, sparking the deep learning boom. in 2012 alongside Alex KrizhevskyAlex KrizhevskyBuilt AlexNet in 2012 with Ilya Sutskever and Geoffrey Hinton, the deep convolutional network whose ImageNet win is widely credited with kicking off the deep learning boom. — the result that restarted deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones. as the field's dominant paradigm. Two years later, with Oriol Vinyals and Quoc Le, he co-authored the sequence-to-sequence learning paper that showed a single neural networkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. could map an input sequence directly to an output sequence — the architecture that made neural machine translation practical and a direct conceptual ancestor of how modern LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. handle variable- length input and output.
In 2015 Sutskever co-founded OpenAI and became its chief scientist, where he was among the strongest internal advocates for the bet that scaling up transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.-based language models — more data, more parameters, more compute — would keep producing qualitatively better capabilities rather than diminishing returns. That bet is what produced the GPT series and, later, RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward.-tuned ChatGPT, the release most credited with moving LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. into mainstream use.
Sutskever left OpenAI in 2024 to found Safe Superintelligence Inc., a lab focused specifically on alignmentAlignmentAlignment is whether a trained model's actual objective and behavior match what its designers intended, split into outer and inner alignment. research ahead of capability research — a framing that puts him, alongside Yoshua BengioYoshua BengioHelped lay the statistical foundations of modern language modeling and co-founded the Montreal deep learning school — now a leading voice on AI safety., among the more safety-focused voices to emerge from within the labs pursuing frontier scale.
See also: Geoffrey HintonGeoffrey HintonKnown as a godfather of deep learning — co-authored backpropagation in 1986, co-invented the Boltzmann machine and dropout, and co-authored AlexNet, the 2012 result that restarted the field., Ashish VaswaniAshish VaswaniLead author of "Attention Is All You Need" (2017), the paper that introduced the transformer architecture behind every modern LLM.
Learn more: LLMs · Wikipedia: Ilya Sutskever
Mentioned in
Lessons where this comes up in context.
Ian Goodfellow
Invented generative adversarial networks (GANs) in 2014, the architecture that first made realistic image generation possible by pitting two networks against each other.
Ashish Vaswani
Lead author of "Attention Is All You Need" (2017), the paper that introduced the transformer architecture behind every modern LLM.