Tokenization
Tokenization converts raw text into a sequence of integers a model can process, typically via subword schemes like byte-pair encoding (BPE).
Tokenization converts raw text into a sequence of integer tokens a neural networkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation. can process. Modern LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. use subword tokenization (commonly byte-pair encoding, BPE): common words get a single token, rarer words split into smaller frequent pieces — keeping sequences short while still representing any string, including unseen words. Each token is then mapped to a learned embeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space. vector before entering the model.
How it works
BPE is trained, not designed. Starting from raw bytes, it repeatedly
counts adjacent symbol pairs across a pretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability.
corpus, merges the most frequent pair into a new symbol, and records
the merge. After N merges you have a vocabulary of N entries plus
the base bytes — typically 32k to 200k in total — and an ordered merge
list. Encoding replays those merges greedily over the input; decoding
is a table lookup and concatenation.
Leading whitespace is usually part of the token, so "cat" and
" cat" are different ids. Because the base alphabet is bytes, every
input encodes — there is no out-of-vocabulary failure, only worse
compression. Special tokens for turn boundaries and end-of-sequence are
appended to the vocabulary rather than produced by merging.
When it breaks
- Character-level tasks. Counting letters, reversing strings, and rhyming are hard because the model never sees characters, only opaque ids.
- Arithmetic and rare names. Digit grouping varies by tokenizer, so
1234may be one token or three, which makes numeric reasoning brittle in ways that shift between models. - Cost is not uniform. Non-English text and code often need several times more tokens per character, inflating inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. latency and price for the same content.
- Mismatch is silent. Serving weights with a different tokenizer or chat template than training used produces fluent nonsense rather than an error — verify them together after any fine-tuneFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions..
See also: EmbeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space., LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama.
Learn more: LLMs · Wikipedia: Byte pair encoding
Mentioned in
Lessons where this comes up in context.
LLM (Large Language Model)
An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama.
Embedding
An embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space.