Breaking Language into Pieces
A transformer eats a sequence of tokens. A token is not a word and not a letter — it is a chunk, usually a few characters. How a text is chopped up is called tokenization, and it is more consequential than it looks.
Byte-pair encoding
Byte-pair encoding (BPE) starts from individual bytes and repeatedly merges the most frequent adjacent pair into a new symbol. Common words become one token; rare words become several. This gives a fixed-size vocabulary that can spell anything, including names and typos, without ever hitting an unknown symbol.
Why you should care
Tokens explain a lot of surprising model behaviour. Prices, dates, and arithmetic are hard partly because numbers split into awkward chunks. Prompts are billed and limited in tokens, not words. Context windows are measured in tokens. And when a model “misses” a letter in a word, it is often because it never saw that letter separately — it saw one token.
The hidden interface
Tokenization is a lossy pre-processing step baked in before training. Changing it later means retraining the model. It is the part of the stack most people ignore and most bugs hide in.