AI Foundationspredict · compress · act
act VII

The Attention Revolution

Pretraining is self-supervised: no labels needed, because every text is its own answer key.

26

The Big Read

before this →Breaking Language into Pieces

Where do you get labels for all the text ever written? You do not need any. Hide the next token and ask the model to predict it. The data labels itself. This is self-supervised learning, and it is the engine of modern AI.

Thecatsatonthe???predict the next tokenmatloss =cross-entropyrepeat for every position in every document
next-token prediction: the answer is always already in the text

One objective, endless data

The task is trivial to state and profound in effect. Given everything so far, produce a probability distribution over the next token, and minimise the cross-entropy against what actually came next. Because every position in every document is a training example, the dataset is effectively unbounded.

To predict a word well, the model is forced to learn syntax, facts, style, reasoning patterns, and something like a world model. It was never told any of that. It is the purest form of the compression thesis: the only pressure is be less surprised.

From raw text to an assistant

A raw pretrained model is a text continuation engine. It has read everything and been instructed in nothing. Ask it a question and it may well answer with more questions, because that is what transcripts look like. Turning it into something that follows instructions, refuses harmful requests, and holds a conversation is the job of the next two chapters.

The three-stage recipe

pretrain → finetune → align

The modern pipeline is now standard. Pretrain on a huge general corpus. Fine-tune on curated examples. Align with human preferences. Each stage is cheaper and smaller than the last, and each changes the model’s character more visibly.

the economyPretraining costs millions and happens once. Alignment costs thousands and happens many times. Most of the product engineering lives in the cheap stages.
PRETRAINpredict text — learn the world
FINE-TUNElearn the task and format
ALIGNlearn our preferences
introduces →pretrainingself-supervised learninglanguage modelnext-token prediction
← previousBreaking Language into Piecesnext →More Is Different