AI Foundationspredict · compress · act
act VII

The Attention Revolution

Loss falls in a straight line on a log-log plot. That simple fact became the strategy of an industry.

27

More Is Different

before this →The Big Read
2020Scaling laws — Loss falls as a smooth power law in compute, data and parameters — so you can plan ahead.
2022Chinchilla — Most giant models were badly under-trained on data. Fix that, and smaller wins.

Plot a model’s loss against its compute on log axes, and you get a startlingly straight line. Loss falls predictably as you add parameters, data, and compute. This is a scaling law, and it turned AI from alchemy into a business plan.

loss (log)compute (log) →measuredfit: a power lawyou can forecast the next model before training it
loss vs. compute on a log-log plot — a straight line you can extrapolate

Why a straight line changes everything

A scaling law lets you plan. Train a series of small models, fit the curve, and predict how well a model ten times larger will do — before you spend the money. Capital allocation became an engineering calculation. This is the direct cause of the multi-hundred-million-dollar training runs of the last decade.

The Chinchilla correction

The first scaling laws over-invested in parameters and under-invested in data. The Chinchilla work (2022) showed that for a fixed compute budget, models were badly undertrained: you should scale data and parameters roughly in proportion. The field had been building giants on too little reading. After Chinchilla, models got smaller and read far more.

Emergence — and the debate

Emergent abilities are capabilities that appear suddenly past a scale threshold: arithmetic, multi-step reasoning, in-context learning. Some researchers argue these are real phase changes. Others argue they are artifacts of how we measure — a smooth capability crossing a sharp test threshold. The disagreement is unresolved, and it matters for safety: if abilities appear abruptly, you may not see the next one coming.

The bitter lesson, restated

Sutton, 2019

Rich Sutton’s essay argues that throughout AI history, general methods that scale with compute have beaten methods that encode human knowledge. Search and learning win; hand-crafted cleverness loses.

the tensionThis is a real tension with chapter 11. Inductive bias helps when data is scarce. When data and compute are abundant, generality wins. Both are true — the question is which regime you are in.
COMPUTEspend more FLOPs
DATAread more tokens — Chinchilla balance
CAPABILITYloss falls; abilities appear
introduces →scaling lawcompute-optimalemergent abilityChinchilla
← previousThe Big Readnext →Teaching Taste