AI Foundationspredict · compress · act

concepts → Benchmark

Benchmark

Benchmark

frontieras of 2026-094 connections

a fixed task suite and metric — the yardstick, and the thing people overfit to

Where it sits

The prism has six jobs across and eight layers down. Its primary cell ismeasure × L6. Hatched cells cannot exist — a GPU does not learn, an institution does not infer.

Overlays: Evaluation & methodology.

What must come first — and what it unlocks

Left to right is reading order, derived from the prerequisite_of edges. Nothing here is hand-ordered: the diagram is the graph.

Turing testBenchmarkEvaluation

Before it: Turing test

Unlocks: Evaluation

How it goes wrong — and what answers that

known failure modes
Contaminationtest questions that leaked into training, making the score a memory test

Where to read it

The chapter that introduces it, and any chapter that uses it again.

32How Do We Even Know?act 7 · The Alignment Layer

Where it comes from

Every connection

All 4 edges touching this node, grouped by relation family — the sections above are highlights from this list. Colours match the relation families inthe atlas.

Order · 2
must come beforeEvaluationField
requiresTuring testThoughtExperiment
Flow · 1
measuresGeneralizationProperty
Failure · 1
fails byContaminationFailureMode

This page is a projection of one node in src/data/concepts.ts. It has no prose file of its own — 230 declared edges produce all 236 of these pages. Edit an edge and both endpoints change.