Modern Deep Learning & Transformers

Noam Shazeer

Co-authored "Attention Is All You Need" and pioneered sparse mixture-of-experts models, then left Google to found Character.AI before returning to lead work on Gemini.

Noam Shazeer was already an unusually prolific Google researcher before co-authoring "Attention Is All You Need" in 2017 alongside Ashish Vaswani and six others — the paper that introduced the transformer. Earlier that same year, Shazeer had led "Outrageously Large Neural Networks," which introduced a sparsely-gated mixture-of-experts layer: instead of running every input through the network's full set of parameters, a learned router sends each input to only a handful of specialized "expert" subnetworks. That idea — spend more total parameters without spending more compute per input — sat mostly dormant for years before becoming a standard technique in frontier-scale LLMs, where mixture-of-experts architectures let labs grow parameter counts without a proportional rise in inference cost.

In 2021, with fellow Google researcher Daniel De Freitas, Shazeer left to found Character.AI, built around large language models fine-tuned for open-ended roleplay and companionship — a distinctly different bet on what people would want from an LLM than the assistant-style products most labs were shipping. In 2024, Google struck a deal to license Character.AI's technology and rehire Shazeer and De Freitas, bringing him back to work on the Gemini model family, alongside Demis Hassabis's Google DeepMind.

See also: Ashish Vaswani, Demis Hassabis

Learn more: Attention & Transformers · Wikipedia: Noam Shazeer