AI Foundationspredict · compress · act
act VII

The Attention Revolution

Models do not see words or letters. They see tokens — and the choice of tokenizer shapes everything downstream.

25

Breaking Language into Pieces

before this →Attention Is All You Need

A transformer eats a sequence of tokens. A token is not a word and not a letter — it is a chunk, usually a few characters. How a text is chopped up is called tokenization, and it is more consequential than it looks.

low·low·lowerstart: characterslolow·w·lowerlowlowerafter merges: common chunks become single tokensthe vocabulary isa fixed list of chunks~50k–200k entrieslearned from data,not designed by hand
byte-pair encoding merges the most frequent pair, over and over

Byte-pair encoding

Byte-pair encoding (BPE) starts from individual bytes and repeatedly merges the most frequent adjacent pair into a new symbol. Common words become one token; rare words become several. This gives a fixed-size vocabulary that can spell anything, including names and typos, without ever hitting an unknown symbol.

Why you should care

Tokens explain a lot of surprising model behaviour. Prices, dates, and arithmetic are hard partly because numbers split into awkward chunks. Prompts are billed and limited in tokens, not words. Context windows are measured in tokens. And when a model “misses” a letter in a word, it is often because it never saw that letter separately — it saw one token.

The hidden interface

every model has one, few users see it

Tokenization is a lossy pre-processing step baked in before training. Changing it later means retraining the model. It is the part of the stack most people ignore and most bugs hide in.

the practical noteWhen something behaves oddly with numbers, code, or rare languages, suspect the tokenizer before the model.
TEXTraw characters in, any language
BPEmerge frequent pairs into chunks
IDSa sequence of integers the model can embed
introduces →tokentokenizerbyte-pair encodingvocabulary
← previousAttention Is All You Neednext →The Big Read