The Big Read
Where do you get labels for all the text ever written? You do not need any. Hide the next token and ask the model to predict it. The data labels itself. This is self-supervised learning, and it is the engine of modern AI.
One objective, endless data
The task is trivial to state and profound in effect. Given everything so far, produce a probability distribution over the next token, and minimise the cross-entropy against what actually came next. Because every position in every document is a training example, the dataset is effectively unbounded.
To predict a word well, the model is forced to learn syntax, facts, style, reasoning patterns, and something like a world model. It was never told any of that. It is the purest form of the compression thesis: the only pressure is be less surprised.
From raw text to an assistant
A raw pretrained model is a text continuation engine. It has read everything and been instructed in nothing. Ask it a question and it may well answer with more questions, because that is what transcripts look like. Turning it into something that follows instructions, refuses harmful requests, and holds a conversation is the job of the next two chapters.
The three-stage recipe
The modern pipeline is now standard. Pretrain on a huge general corpus. Fine-tune on curated examples. Align with human preferences. Each stage is cheaper and smaller than the last, and each changes the model’s character more visibly.