AI Foundationspredict · compress · act

Every concept

The graph has 236 nodes and 230 declared edges. Each one has a page — its position on the prism, what must come first, what it unlocks, how it fails, and every connection it has. The list is ordered by centrality: how many edges touch it, which is a rough measure of how much of the rest of the field leans on it.

the atlas →by kind →reading routes →

#conceptlinks
01Agent harnessSystemevery decision you can make without retraining the model — loop, context, tools, memory13
02Reinforcement learningFieldlearning what to do by trying, from delayed reward10
03AgentSystema model plus a harness, running a loop toward a goal9
04AttentionMechanismlet each token weigh every other token, instead of only the last one9
05EntropyQuantitythe average surprise of a source, in bits9
06Tool useMechanismthe model asks for a function to be run, and reads the result back9
07Diffusion modelArchitecturelearn to undo noise, and you can make an image from nothing8
08Model familyPropertyone line of models at several sizes, sharing a recipe and a name8
09PretrainingAlgorithmlearn from raw text at scale by predicting what comes next8
10RLHFAlgorithmtilt a pretrained model toward human preference — how raw prediction becomes behaviour7
11TransformerArchitectureattention and feed-forward blocks stacked — the architecture that swallowed the field7
12GeneralizationPropertydoing well on data you have never seen — the only thing that counts6
13Word embeddingMechanismwords as coordinates, so that similar meanings sit near each other6
14Artificial neuronMechanismadd the inputs, then pass the total through a simple curve5
15Bayes' ruleAlgorithmturn a prior and a likelihood into a posterior — updating belief with evidence5
16Fine-tuningAlgorithmcontinue training a pretrained model on a smaller, targeted set5
17InferenceMechanismrunning a trained model to get an answer out5
18Language modelArchitecturea model that puts a probability on the next piece of text5
19Model Context ProtocolProtocola standard way to plug tools into a model, so every model need not learn every tool5
20Retrieval-augmented generationAlgorithmlook facts up and put them in the context, instead of storing them in the weights5
21Scaling lawResultloss falls as a smooth power law in compute, data and parameters — so you can plan ahead5
22TokenizerMechanismthe rule that cuts text into tokens — it shapes everything downstream5
23BackpropagationAlgorithmassign blame for the error backwards through the layers, then nudge every weight4
24BenchmarkBenchmarka fixed task suite and metric — the yardstick, and the thing people overfit to4
25Convolutional neural networkArchitecturestacked convolutions that see images through learned local filters4
26Cross-entropyObjectivethe cost of coding data with the wrong model — the loss nearly everything trains on4
27DeepSeek-V3Modela 671B MoE (37B active) trained in FP8 — frontier-adjacent at a fraction of the assumed cost4
28Mixture of expertsMechanismroute each token to a few small experts — huge capacity, small active cost4
29Multilayer perceptronArchitecturestacked neuron layers that can bend a boundary no single perceptron could4
30OrchestrationSystemthe layer that decides what runs, in what order, and what happens when it fails4
31PerceptronMechanismthe first machine that learned weights to separate two classes4
32Reasoning modelArchitecturea model trained to spend more thinking before it answers4
33Recurrent neural networkArchitecturea network that carries a hidden state forward, one step at a time4
34Agent loopMechanismthink, act, look at the result, repeat until the job is done3
35ComputabilityPropertywhether some step-by-step method can settle a question at all3
36Compute-optimalPropertythe best split of a fixed compute budget between model size and data3
37Control theoryFieldthe mathematics of steering a system toward a target without it wobbling3
38EvaluationFielddeciding what a number means — validity, contamination, ablations3
39GPUHardwarethousands of simple cores doing the same maths at once — the accident that trains everything3
40Gradient descentAlgorithmwalk downhill on the loss by following its slope3
41GRPOAlgorithmscore a group of answers against each other, so no separate value network is needed3
42InterpretabilityFieldfinding out what the insides are actually doing3
43Next-token predictionTaskthe one training game behind a language model: guess the next token3
44ObservabilityFieldbeing able to see what a running system is doing and why it failed3
45Omni-modelArchitectureone model for text, audio, images and video, instead of a relay of specialists3
46ProbabilityFieldthe mathematics of belief under uncertainty3
47QuantizationMechanismstore weights in fewer bits, so a model fits on hardware you own3
48Symbolic AIFieldintelligence as rules over written symbols, added by hand3
49World modelArchitecturean internal simulator of how a situation evolves — imagining without acting3
50AlignmentFieldmaking a system do what was meant, not merely what was written2
51Any-to-anyPropertyone model that takes in text, image, audio or video and emits any of them2
52Backpropagation through timeAlgorithmunroll a loop over time, then run backprop through the whole tape2
53CapacityQuantityhow many different patterns a model is able to fit2
54ChannelMechanismthe pipe a message travels through, which may corrupt it2
55ChinchillaResultfor a fixed budget, most big models of 2020 were too large and under-trained2
56CompressionHeuristicintelligence is compression — predicting well and compressing well are one thing2
57ConvolutionMechanismslide one small filter across the input and reuse it everywhere — built-in translation symmetry2
58DeepSeek-R1Modelopen reasoning weights under MIT — R1-Zero learned to reason with no supervised examples2
59Diffusion transformerArchitecturea transformer as the image denoiser, usually trained with flow matching2
60Durable executionMechanisma long run that survives crashes and restarts where it left off2
61Embedding spacePhenomenonthe geometry where similar things end up near each other2
62Hidden layerMechanisma layer of neurons between input and output that builds its own features2
63Huawei AscendHardwarea non-Nvidia AI accelerator line, pushed hard by export controls2
64Inference scalingPropertytrading more thinking at answer time for better answers2
65InformationQuantityhow much uncertainty a message removes2
66KL divergenceQuantityextra bits wasted by believing q when the truth is p — never negative, never symmetric2
67Knowledge representationMechanismwriting what the system knows in a form a machine can reason over2
68Latent diffusionArchitecturerun the diffusion steps in a small squeezed space, then decode back to pixels2
69Learning rateQuantityhow big a step to take downhill — too small is slow, too big overshoots2
70LikelihoodQuantityhow probable the data you saw would be under a given hypothesis2
71Loss functionObjectivethe number training tries to make small — the formal statement of what you want2
72LSTMArchitecturea gated loop that remembers over long stretches and eases the vanishing gradient2
73Mechanistic interpretabilityFieldtaking the model apart to name the parts and what each one does2
74Model-based RLMechanismlearn a model of the world, then plan inside it instead of acting blindly2
75Mutual informationQuantityhow much knowing one thing tells you about another2
76Open weightsPropertythe numbers released, letting anyone run and fine-tune — but not rebuild2
77OverfittingFailureModememorising the training set instead of learning the pattern2
78PermissionMechanismthe grant that lets a program read this file or call that service — and nothing else2
79PlanningMechanismchoosing a whole sequence of steps in advance of doing them2
80Prefix codeMechanisma code where no codeword is the start of another, so it is readable as it arrives2
81PriorDistributionwhat you believed before seeing the evidence2
82RateQuantitythe number of bits used per symbol — the score compression tries to lower2
83RegularizationMechanismany penalty that keeps a model simple so it stops memorising2
84RewardObjectivea single number saying how well things just went2
85SandboxMechanisman isolated place for an agent to act, so a wrong step is survivable2
86SearchAlgorithmtry possibilities in a smart order instead of all of them2
87Self-supervised learningMechanismhide part of the data and let the model teach itself to recover it2
88Sequence modelArchitectureany model built to read or write things in order2
89Source coding theoremTheoremyou cannot compress below the entropy — and you can get arbitrarily close2
90TokenQuantitya chunk of text — often a word-piece — that the model reads as one unit2
91Training setDatasetthe examples the model is allowed to learn from2
92Turing testThoughtExperimentcan you tell the machine from a person in conversation? — a test of behaviour, not mind2
93Vector databaseSystema store that finds the nearest vectors fast, even among billions2
94Activation functionMechanismthe bend a neuron applies to its weighted sum — without it, layers collapse into one1
95APIProtocola fixed doorway one program uses to ask another for something1
96Auxiliary-loss-free load balancingMechanismkeep experts evenly used by nudging a bias, instead of adding a penalty to the loss1
97BatchingMechanismserving many requests together so the hardware is never idle1
98Bayesian inferenceAlgorithmreasoning that keeps a full distribution over hypotheses and updates it1
99BiasQuantitya neuron's resting offset — how eager it is to fire with no input1
100Bias–variance tradeoffPropertyerror splits into being too simple and being too jumpy; lowering one raises the other1
101BitQuantityone yes-or-no of uncertainty — the atom of information1
102Byte-pair encodingAlgorithmbuild a vocabulary by repeatedly gluing together the most common pair of pieces1
103CANNSystemHuawei's software stack for Ascend chips — the counterpart to CUDA1
104Chain of thoughtMechanismwriting the working out, which gives the model more room to be right1
105Channel capacityTheoremthe most information a noisy channel can carry — a hard ceiling1
106CheckpointingMechanismsaving progress at safe points so a failure does not undo everything1
107ChunkingMechanismcutting long documents into pieces small enough to retrieve and fit1
108Classifier-free guidanceMechanismsteer generation harder toward the prompt by contrasting with and without it1
109Code executionMechanismletting the model write and run real code instead of guessing the answer1
110Combinatorial explosionFailureModeoptions multiply so fast that brute force stops being possible1
111Computer useMechanismdriving a real screen with mouse and keyboard, like a person would1
112ConstitutionDocumenta written set of principles a model is trained to critique itself against1
113ContaminationFailureModetest questions that leaked into training, making the score a memory test1
114Context managementMechanismdeciding what to keep, summarise or drop so the model sees the right things1
115Context windowPropertythe most text a model can hold in mind at once1
116ControlNetArchitecturea side branch that lets a sketch or pose steer an image model without retraining it1
117CyberneticsFieldthe 1940s study of control and communication in animals and machines alike1
118DeepSeek-V4Modelthe 2026 follow-up — 1.6T/49B and 284B/13B MoE, one-million-token context, hybrid attention replacing MLA1
119DenoisingMechanismlearning to remove a little noise, so many tiny steps turn noise into an image1
120Direct preference optimisationAlgorithmlearn straight from preference pairs, skipping the separate reward model1
121Distributional hypothesisHeuristicwords used in similar company end up meaning similar things1
122DivergencePropertya measure of how far two distributions are apart, without being a distance1
123Dot productQuantitymultiply matching entries and add — a number for how aligned two vectors are1
124DriftPhenomenonthe world shifts, so a model that was right slowly stops being right1
125Emergent abilityPhenomenona skill that barely shows, then appears suddenly as the model grows1
126EnvironmentSystemeverything outside the agent that it can sense and change1
127Error-correcting codeAlgorithmadd redundancy so the receiver can undo the channel's damage1
128Evaluation harnessSystemthe machinery that runs a suite of tests and scores the results the same way every time1
129Expert systemSystema hand-built rulebook that answers like a specialist in a narrow domain1
130ExplorationMechanismdeliberately trying something new so you learn what you are missing1
131Export controlsDocumentrules limiting chip sales across borders — the constraint that bent the field toward efficiency, then partly loosened1
132FeatureVariablea direction inside the model that stands for one human-readable idea1
133Feature mapQuantitythe grid of activations a filter produces as it slides across an image1
134FeedbackMechanismmeasure the gap to the goal and use it to correct the next step1
135Flow matchingMechanismlearn a straight path from noise to data instead of a slow curve1
136FP8Formatan 8-bit number format for training and serving, roughly half the memory of 16-bit1
137Frame problemFailureModehard to say which facts an action leaves unchanged, without listing them all1
138Fully open modelPropertyopen weights plus open data, code and logs — reproducible end to end1
139GGUFFormata single self-contained file holding quantised weights and their metadata1
140GLMModelZhipu's model family, from GLM-130B through the GLM-4 line1
141GuardrailMechanisma check that catches and blocks a bad output before a user sees it1
142Halting problemResultno program can decide, for every program, whether it will ever stop1
143Harness engineeringFielddesigning the loop, tools, memory and permissions around a model — now a discipline of its own1
144Huffman codingAlgorithmbuild the optimal prefix code by repeatedly merging the two least likely symbols1
145Human in the loopMechanisma person approves the risky step before the agent takes it1
146HunyuanModelTencent's open mixture-of-experts line, hundreds of billions of parameters1
147Inductive biasPropertythe assumptions a learner makes so it can guess about unseen cases1
148Instruction tuningMechanismfine-tune on examples written as tasks, so following requests becomes the habit1
149IntelligencePropertyreaching goals in an environment, across many different situations1
150InterconnectMechanismthe links between chips — often the real limit on training a large model1
151InternLMModelShanghai AI Laboratory's open model line, released with full technical reports1
152InteroperabilityPropertytools and agents from different makers working together without custom glue1
153JEPAArchitecturepredict in a learned abstract space, not pixel by pixel1
154KimiModelMoonshot's long-context line, with reasoning trained by large-scale RL1
155Knowledge distillationMechanismtrain a small model to copy a big one's answers, gaining skill cheaply1
156KV cacheMechanismstore the keys and values you already computed, so each new token is cheap1
157Latent spacePropertya squeezed code where each point stands for a whole image, far smaller than pixels1
158Least privilegeHeuristicgive every part the smallest power it needs to do the job1
159LlamaModelthe family that made open weights mainstream — under a community licence with a user cap1
160Local servingSystemrunning a model on your own machine, with no network round-trip1
161Loop policyHeuristicthe rules for when to keep going, ask a human, or stop1
162LoRAMechanismfreeze the model and train a tiny add-on — fine-tuning on one GPU1
163Markov decision processMechanismstates, actions and rewards, where the next state depends only on now1
164MatrixQuantitya grid of numbers acting as one machine that turns vectors into vectors1
165MCP clientSystemthe agent side that connects to servers and offers their tools to the model1
166MCP serverSystema small program that exposes tools and data through the shared protocol1
167Memory bandwidthQuantityhow fast numbers move from memory — the limit that decoding usually hits first1
168Memory wallPhenomenoncompute got fast but memory did not, so moving numbers is now the bottleneck1
169MiniMaxModela Chinese lab building long-context models on lightning attention and mixture of experts1
170MistralModela European lab whose small open models pushed quality per parameter1
171Modality encoderArchitecturethe part that turns one kind of input — sound, image, video — into shared tokens1
172Multi-agentSystemseveral agents splitting a job, talking to each other as they go1
173Multi-head latent attentionMechanismcompress the per-token memory attention must keep, so long contexts cost far less1
174Multi-token predictionMechanismtrain the model to guess several next tokens at once, not just one1
175NoisePhenomenonrandom damage the channel adds to the message1
176Occam's razorHeuristicprefer the simplest explanation that fits — a guess, not a theorem1
177OLMoModela fully open model line that releases data, code and training logs too1
178Open image modelModelan image generator you can download and run yourself1
179Open-weight licenceDocumentthe terms attached to downloadable weights — open weights are not always open source1
180ParallelismPropertydoing many small sums at once instead of one after another1
181PerplexityQuantityhow many equally likely choices the model felt it was facing — cross-entropy, exponentiated1
182PolicyMechanismthe rule that says which action to take in each situation1
183PoolingMechanismshrink a feature map by keeping the strongest value in each patch1
184Positional encodingMechanisma pattern added so the model knows where each token sat in the order1
185PosteriorDistributionwhat you believe after seeing the evidence1
186PredictionMechanismsaying what comes next before you are told1
187Prompt injectionFailureModeinstructions hidden in content the model reads, obeyed as if you had typed them1
188Quantised weightsFormatweights stored at 8, 4 or fewer bits so a big model fits in small memory1
189Query, key, valueMechanismeach token asks a question, every token advertises what it holds, and answers are blended by fit1
190QwenModelAlibaba's open model line, released across many sizes and modalities1
191ReActAlgorithminterleave a line of reasoning with a tool call, so thought follows evidence1
192Red teamingMechanismattacking the system on purpose to find how it breaks before others do1
193Reward hackingFailureModegetting the reward without doing the thing — the optimiser finds the loophole1
194Reward modelModela model trained on human preferences that scores answers so training can use it1
195Self-supervised videoMechanismlearn from raw video by predicting what happens next in it1
196Semantic searchMechanismfinding things by meaning rather than by matching words1
197Sparse autoencoderArchitecturea tool that pulls many tidy, mostly-silent features out of a crowded layer1
198SpecificationDocumentwriting down what we actually want the system to do1
199Speculative decodingAlgorithma small fast model drafts, a big model checks — same output, fewer big steps1
200Speech tokenQuantitya chunk of sound turned into something the same machinery can treat like a word1
201SuperpositionPhenomenona model packs more ideas than it has directions, by sharing them at angles1
202SurpriseQuantityhow unlikely the thing that happened was — the raw material of entropy1
203Tensor coreHardwarea circuit that does a whole small matrix multiply in one instruction1
204Test setDatasetexamples held back, used only to check whether learning generalised1
205Test-time computeQuantitythe effort spent while answering, not while training1
206Tool schemaFormata written description of a function — its name, inputs and outputs — a model can be handed1
207TracingMechanismrecording every step, input and output so a run can be replayed1
208Turing machineMechanisman imagined tape-reading machine that pins down what "computable" means1
209UncertaintyPropertyhow many answers are still possible1
210UnderfittingFailureModethe model is too simple to capture the pattern, so it is wrong everywhere1
211Value functionEstimatora guess of how much reward is still to come from here1
212Vanishing gradientFailureModein a deep stack the learning signal fades to nothing before it reaches the bottom1
213VectorQuantitya list of numbers treated as one point or one arrow1
214Vector arithmeticPhenomenonking − man + woman ≈ queen: directions in embedding space carry meaning1
215vLLMSystema serving engine whose paged memory made long-context batching practical1
216VocabularyDatasetthe fixed set of tokens a model can read and write1
217WeightQuantityhow much a neuron cares about one input — the number training changes1
218word2vecAlgorithmlearn a vector per word by predicting its neighbours, cheaply and at scale1
219WorkflowMechanisma fixed sequence of steps, with the model doing only the fuzzy parts1
220Acceptable use policyDocumenta list of things you may not do with a model, attached to its licence0
221AGIOpenQuestiona system that can do any cognitive job a person can — nobody agrees on the test0
222AI winterPhenomenona period when promises outran results and the money left0
223BarbellPhenomenonthe field splits toward a few huge closed models and many small open ones0
224CommoditisationPhenomenoncapability that was rare becomes cheap, so value moves to what you build on top0
225EU AI ActDocumentthe European Union law that sorts AI uses into risk tiers and sets duties for each0
226Full-duplexPropertylisten and speak at the same time, so talking feels like a conversation0
227GPAIPropertygeneral-purpose AI — the EU category for large models usable for many tasks0
228Model cardDocumenta short factual page describing what a model is, how it was made and its limits0
229No free lunchTheoremaveraged over all problems, no learner beats any other — so assumptions are everything0
230Open-source AIFieldbuilding AI in the open, where anyone can read, run and change it0
231Open-weight baselineBenchmarkthe best open model, used as the yardstick everyone else compares against0
232Open-weight riskOpenQuestiondoes releasing weights let the dangerous uses scale faster than the defensive ones?0
233Openness ladderPropertya scale from weights-only to fully open data, code and logs0
234Permissive licenceDocumenta licence that lets you use, change and sell with almost no conditions0
235SuperintelligenceOpenQuestiona system smarter than us in every way — a hypothesis, not an observation0
236The scaling debateOpenQuestionis more scale enough for general capability, or is something still missing?0

17 nodes carry no edges yet — they are in the graph ahead of their connections. Nothing on this page is hand-listed: it is the node array, sorted by degree.