Modern Deep Learning & Transformers

Ashish Vaswani

Lead author of "Attention Is All You Need" (2017), the paper that introduced the transformer architecture behind every modern LLM.

Ashish Vaswani led the eight-author team at Google that published "Attention Is All You Need" in 2017 — the paper that introduced the transformer. Its central claim was that attention alone, without the recurrence used by RNNs or the local receptive fields used by CNNs, was sufficient to model sequences, and that removing recurrence's inherently sequential computation meant every position in a sequence could be processed in parallel during training. That single architectural change is why training runs that would have taken months on an RNN became feasible in days on a transformer, and it's the reason every major LLM built since — GPT, BERT, Claude, and the rest — shares the same underlying architecture.

The paper's co-authors — Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, and Illia Polosukhin — have themselves gone on to found or shape several major AI labs and products, making "Attention Is All You Need" one of the most consequential single papers in the field's history, cited well over 100,000 times.

Vaswani later co-founded Adept AI and Essential AI, both focused on applying transformer-based models to autonomous agents and enterprise tasks — a natural extension of an architecture originally designed for machine translation.

See also: Ilya Sutskever, Noam Shazeer

Learn more: Attention & Transformers · Wikipedia: Attention Is All You Need