AI Foundationspredict · compress · act

concepts → Result

Result

Scaling law

decade5 connections

loss falls as a smooth power law in compute, data and parameters — so you can plan ahead

Where it sits

The prism has six jobs across and eight layers down. Its primary cell ismeasure × L4, with a noted second cell. Hatched cells cannot exist — a GPU does not learn, an institution does not infer.

Overlays: Efficiency, cost & energy. learn×L4 — it dictates how training runs are budgeted

What must come first — and what it unlocks

Left to right is reading order, derived from the prerequisite_of edges. Nothing here is hand-ordered: the diagram is the graph.

nothing comes first — this is a starting pointScaling lawCompute-optimal

Unlocks: Compute-optimal

What kind of thing it is — and what it is made of

It is made of: Compute-optimal, Emergent ability.

Where to read it

The chapter that introduces it, and any chapter that uses it again.

27More Is Differentact 6 · The Attention Revolution

Where it comes from

paperScaling Laws for Neural Language ModelsJared Kaplan, et al. · 2020
paperTraining Compute-Optimal Large Language ModelsJordan Hoffmann, et al. · 2022

Every connection

All 5 edges touching this node, grouped by relation family — the sections above are highlights from this list. Colours match the relation families inthe atlas.

Structure · 2
has as a partCompute-optimalProperty
has as a partEmergent abilityPhenomenon
Order · 1
must come beforeCompute-optimalProperty
Lineage · 1
is improved byChinchillaResult
Contrast · 1
trades off withInference scalingProperty

This page is a projection of one node in src/data/concepts.ts. It has no prose file of its own — 230 declared edges produce all 236 of these pages. Edit an edge and both endpoints change.