Modern Deep Learning & Transformers

Ilya Sutskever

Co-authored AlexNet as a student, then sequence-to-sequence learning, then co-founded OpenAI and helped drive the bet that scale would produce GPT-level language models.

Ilya Sutskever was a graduate student under Geoffrey Hinton when he co-authored AlexNet in 2012 alongside Alex Krizhevsky — the result that restarted deep learning as the field's dominant paradigm. Two years later, with Oriol Vinyals and Quoc Le, he co-authored the sequence-to-sequence learning paper that showed a single neural network could map an input sequence directly to an output sequence — the architecture that made neural machine translation practical and a direct conceptual ancestor of how modern LLMs handle variable- length input and output.

In 2015 Sutskever co-founded OpenAI and became its chief scientist, where he was among the strongest internal advocates for the bet that scaling up transformer-based language models — more data, more parameters, more compute — would keep producing qualitatively better capabilities rather than diminishing returns. That bet is what produced the GPT series and, later, RLHF-tuned ChatGPT, the release most credited with moving LLMs into mainstream use.

Sutskever left OpenAI in 2024 to found Safe Superintelligence Inc., a lab focused specifically on alignment research ahead of capability research — a framing that puts him, alongside Yoshua Bengio, among the more safety-focused voices to emerge from within the labs pursuing frontier scale.

See also: Geoffrey Hinton, Ashish Vaswani

Learn more: LLMs · Wikipedia: Ilya Sutskever

Mentioned in

Lessons where this comes up in context.