Every concept
The graph has 236 nodes and 230 declared edges. Each one has a page — its position on the prism, what must come first, what it unlocks, how it fails, and every connection it has. The list is ordered by centrality: how many edges touch it, which is a rough measure of how much of the rest of the field leans on it.
the atlas →by kind →reading routes →
#conceptlinks
01Agent harnessSystemevery decision you can make without retraining the model — loop, context, tools, memory13
10RLHFAlgorithmtilt a pretrained model toward human preference — how raw prediction becomes behaviour7
11TransformerArchitectureattention and feed-forward blocks stacked — the architecture that swallowed the field7
15Bayes' ruleAlgorithmturn a prior and a likelihood into a posterior — updating belief with evidence5
19Model Context ProtocolProtocola standard way to plug tools into a model, so every model need not learn every tool5
20Retrieval-augmented generationAlgorithmlook facts up and put them in the context, instead of storing them in the weights5
21Scaling lawResultloss falls as a smooth power law in compute, data and parameters — so you can plan ahead5
23BackpropagationAlgorithmassign blame for the error backwards through the layers, then nudge every weight4
25Convolutional neural networkArchitecturestacked convolutions that see images through learned local filters4
26Cross-entropyObjectivethe cost of coding data with the wrong model — the loss nearly everything trains on4
27DeepSeek-V3Modela 671B MoE (37B active) trained in FP8 — frontier-adjacent at a fraction of the assumed cost4
28Mixture of expertsMechanismroute each token to a few small experts — huge capacity, small active cost4
29Multilayer perceptronArchitecturestacked neuron layers that can bend a boundary no single perceptron could4
30OrchestrationSystemthe layer that decides what runs, in what order, and what happens when it fails4
33Recurrent neural networkArchitecturea network that carries a hidden state forward, one step at a time4
39GPUHardwarethousands of simple cores doing the same maths at once — the accident that trains everything3
45Omni-modelArchitectureone model for text, audio, images and video, instead of a relay of specialists3
49World modelArchitecturean internal simulator of how a situation evolves — imagining without acting3
52Backpropagation through timeAlgorithmunroll a loop over time, then run backprop through the whole tape2
56CompressionHeuristicintelligence is compression — predicting well and compressing well are one thing2
57ConvolutionMechanismslide one small filter across the input and reuse it everywhere — built-in translation symmetry2
58DeepSeek-R1Modelopen reasoning weights under MIT — R1-Zero learned to reason with no supervised examples2
59Diffusion transformerArchitecturea transformer as the image denoiser, usually trained with flow matching2
66KL divergenceQuantityextra bits wasted by believing q when the truth is p — never negative, never symmetric2
67Knowledge representationMechanismwriting what the system knows in a form a machine can reason over2
68Latent diffusionArchitecturerun the diffusion steps in a small squeezed space, then decode back to pixels2
71Loss functionObjectivethe number training tries to make small — the formal statement of what you want2
78PermissionMechanismthe grant that lets a program read this file or call that service — and nothing else2
80Prefix codeMechanisma code where no codeword is the start of another, so it is readable as it arrives2
87Self-supervised learningMechanismhide part of the data and let the model teach itself to recover it2
89Source coding theoremTheoremyou cannot compress below the entropy — and you can get arbitrarily close2
92Turing testThoughtExperimentcan you tell the machine from a person in conversation? — a test of behaviour, not mind2
94Activation functionMechanismthe bend a neuron applies to its weighted sum — without it, layers collapse into one1
96Auxiliary-loss-free load balancingMechanismkeep experts evenly used by nudging a bias, instead of adding a penalty to the loss1
98Bayesian inferenceAlgorithmreasoning that keeps a full distribution over hypotheses and updates it1
100Bias–variance tradeoffPropertyerror splits into being too simple and being too jumpy; lowering one raises the other1
102Byte-pair encodingAlgorithmbuild a vocabulary by repeatedly gluing together the most common pair of pieces1
108Classifier-free guidanceMechanismsteer generation harder toward the prompt by contrasting with and without it1
110Combinatorial explosionFailureModeoptions multiply so fast that brute force stops being possible1
113ContaminationFailureModetest questions that leaked into training, making the score a memory test1
114Context managementMechanismdeciding what to keep, summarise or drop so the model sees the right things1
116ControlNetArchitecturea side branch that lets a sketch or pose steer an image model without retraining it1
118DeepSeek-V4Modelthe 2026 follow-up — 1.6T/49B and 284B/13B MoE, one-million-token context, hybrid attention replacing MLA1
119DenoisingMechanismlearning to remove a little noise, so many tiny steps turn noise into an image1
120Direct preference optimisationAlgorithmlearn straight from preference pairs, skipping the separate reward model1
128Evaluation harnessSystemthe machinery that runs a suite of tests and scores the results the same way every time1
131Export controlsDocumentrules limiting chip sales across borders — the constraint that bent the field toward efficiency, then partly loosened1
137Frame problemFailureModehard to say which facts an action leaves unchanged, without listing them all1
143Harness engineeringFielddesigning the loop, tools, memory and permissions around a model — now a discipline of its own1
144Huffman codingAlgorithmbuild the optimal prefix code by repeatedly merging the two least likely symbols1
148Instruction tuningMechanismfine-tune on examples written as tasks, so following requests becomes the habit1
152InteroperabilityPropertytools and agents from different makers working together without custom glue1
155Knowledge distillationMechanismtrain a small model to copy a big one's answers, gaining skill cheaply1
157Latent spacePropertya squeezed code where each point stands for a whole image, far smaller than pixels1
159LlamaModelthe family that made open weights mainstream — under a community licence with a user cap1
163Markov decision processMechanismstates, actions and rewards, where the next state depends only on now1
167Memory bandwidthQuantityhow fast numbers move from memory — the limit that decoding usually hits first1
168Memory wallPhenomenoncompute got fast but memory did not, so moving numbers is now the bottleneck1
169MiniMaxModela Chinese lab building long-context models on lightning attention and mixture of experts1
171Modality encoderArchitecturethe part that turns one kind of input — sound, image, video — into shared tokens1
173Multi-head latent attentionMechanismcompress the per-token memory attention must keep, so long contexts cost far less1
174Multi-token predictionMechanismtrain the model to guess several next tokens at once, not just one1
179Open-weight licenceDocumentthe terms attached to downloadable weights — open weights are not always open source1
181PerplexityQuantityhow many equally likely choices the model felt it was facing — cross-entropy, exponentiated1
184Positional encodingMechanisma pattern added so the model knows where each token sat in the order1
187Prompt injectionFailureModeinstructions hidden in content the model reads, obeyed as if you had typed them1
189Query, key, valueMechanismeach token asks a question, every token advertises what it holds, and answers are blended by fit1
193Reward hackingFailureModegetting the reward without doing the thing — the optimiser finds the loophole1
194Reward modelModela model trained on human preferences that scores answers so training can use it1
197Sparse autoencoderArchitecturea tool that pulls many tidy, mostly-silent features out of a crowded layer1
199Speculative decodingAlgorithma small fast model drafts, a big model checks — same output, fewer big steps1
200Speech tokenQuantitya chunk of sound turned into something the same machinery can treat like a word1
201SuperpositionPhenomenona model packs more ideas than it has directions, by sharing them at angles1
206Tool schemaFormata written description of a function — its name, inputs and outputs — a model can be handed1
210UnderfittingFailureModethe model is too simple to capture the pattern, so it is wrong everywhere1
212Vanishing gradientFailureModein a deep stack the learning signal fades to nothing before it reaches the bottom1
214Vector arithmeticPhenomenonking − man + woman ≈ queen: directions in embedding space carry meaning1
220Acceptable use policyDocumenta list of things you may not do with a model, attached to its licence0
224CommoditisationPhenomenoncapability that was rare becomes cheap, so value moves to what you build on top0
225EU AI ActDocumentthe European Union law that sorts AI uses into risk tiers and sets duties for each0
228Model cardDocumenta short factual page describing what a model is, how it was made and its limits0
229No free lunchTheoremaveraged over all problems, no learner beats any other — so assumptions are everything0
231Open-weight baselineBenchmarkthe best open model, used as the yardstick everyone else compares against0
232Open-weight riskOpenQuestiondoes releasing weights let the dangerous uses scale faster than the defensive ones?0
234Permissive licenceDocumenta licence that lets you use, change and sell with almost no conditions0
235SuperintelligenceOpenQuestiona system smarter than us in every way — a hypothesis, not an observation0
236The scaling debateOpenQuestionis more scale enough for general capability, or is something still missing?0
17 nodes carry no edges yet — they are in the graph ahead of their connections. Nothing on this page is hand-listed: it is the node array, sorted by degree.