The Big Read
Where do you get labels for all the text ever written? You do not need any. Hide the next token and ask the model to predict it. The data labels itself. This is self-supervised learning, and it is the engine of modern AI.
One objective, endless data
The task is trivial to state and profound in effect. Given everything so far, produce a probability distribution over the next token, and minimise the cross-entropy against what actually came next. Because every position in every document is a training example, the dataset is effectively unbounded.
To predict a word well, the model is forced to learn syntax, facts, style, reasoning patterns, and something like a world model. It was never told any of that. It is the purest form of the compression thesis: the only pressure is be less surprised.
The data wall
“Effectively unbounded” is not the same as infinite. The high-quality human text on the public web is a finite stock, and at the rate training corpora have grown it may be substantially consumed within the decade — a projection, not a fact, but a serious one. The pressure it creates is already visible: labs license archives, filter and re-filter what they already have, reuse the same tokens for more than one epoch, and generate synthetic text to fill the gap. Each answer costs something. Licensing concentrates data behind paywalls, repetition risks the model laundering its own errors, and synthetic data can drift away from the real distribution it was meant to extend. Running out of text is not a cliff; it is a slow change in what the field’s most basic resource costs.
From raw text to an assistant
A raw pretrained model is a text continuation engine. It has read everything and been instructed in nothing. Ask it a question and it may well answer with more questions, because that is what transcripts look like. Turning it into something that follows instructions, refuses harmful requests, and holds a conversation is the job of the next two chapters.
The three-stage recipe
The modern pipeline is now standard. Pretrain on a huge general corpus. Fine-tune on curated examples. Align with human preferences. Each stage is cheaper and smaller than the last, and each changes the model’s character more visibly.