AI Foundationspredict · compress · act

concepts → Algorithm

Algorithm

Pretraining

decade8 connections

learn from raw text at scale by predicting what comes next

Where it sits

The prism has six jobs across and eight layers down. Its primary cell islearn × L4. Hatched cells cannot exist — a GPU does not learn, an institution does not infer.

What must come first — and what it unlocks

Left to right is reading order, derived from the prerequisite_of edges. Nothing here is hand-ordered: the diagram is the graph.

BackpropagationTransformerTokenizerSelf-supervised learningPretrainingFine-tuning

Before it: Backpropagation · Transformer · Tokenizer · Self-supervised learning

Unlocks: Fine-tuning

What kind of thing it is — and what it is made of

It is made of: Next-token prediction, Multi-token prediction.

Where to read it

The chapter that introduces it, and any chapter that uses it again.

26The Big Readact 6 · The Attention Revolution

Where it comes from

paperLanguage Models are Few-Shot LearnersTom Brown, et al. · 2020

Every connection

All 8 edges touching this node, grouped by relation family — the sections above are highlights from this list. Colours match the relation families inthe atlas.

Structure · 2
has as a partNext-token predictionTask
has as a partMulti-token predictionMechanism
Order · 5
requiresBackpropagationAlgorithm
requiresTransformerArchitecture
requiresTokenizerMechanism
must come beforeFine-tuningAlgorithm
requiresSelf-supervised learningMechanism
Flow · 1
usesCross-entropyObjective

This page is a projection of one node in src/data/concepts.ts. It has no prose file of its own — 230 declared edges produce all 236 of these pages. Edit an edge and both endpoints change.