More Is Different
Plot a model’s loss against its compute on log axes, and you get a startlingly straight line. Loss falls predictably as you add parameters, data, and compute. This is a scaling law, and it turned AI from alchemy into a business plan.
Why a straight line changes everything
A scaling law lets you plan. Train a series of small models, fit the curve, and predict how well a model ten times larger will do — before you spend the money. Capital allocation became an engineering calculation. This is the direct cause of the multi-hundred-million-dollar training runs of the last decade.
The Chinchilla correction
The first scaling laws over-invested in parameters and under-invested in data. The Chinchilla work (2022) showed that for a fixed compute budget, models were badly undertrained: you should scale data and parameters roughly in proportion. The field had been building giants on too little reading. After Chinchilla, models got smaller and read far more.
Emergence — and the debate
Emergent abilities are capabilities that appear suddenly past a scale threshold: arithmetic, multi-step reasoning, in-context learning. Some researchers argue these are real phase changes. Others argue they are artifacts of how we measure — a smooth capability crossing a sharp test threshold. The disagreement is unresolved, and it matters for safety: if abilities appear abruptly, you may not see the next one coming.
The bitter lesson, restated
Rich Sutton’s essay argues that throughout AI history, general methods that scale with compute have beaten methods that encode human knowledge. Search and learning win; hand-crafted cleverness loses.