The glossary
Every term the guide introduces, in alphabetical order. Each is linked to the chapter where it is born. 231 of the 231 are already nodes in theatlas graph and show their kind; the rest are queued to be absorbed into it.
A
acceptable use policyDocumenta list of things you may not do with a model, attached to its licence49Licences and Rules
activation functionMechanismthe bend a neuron applies to its weighted sum — without it, layers collapse into one14The First Neuron
agentSystema model plus a harness, running a loop toward a goal01What Is Intelligence, Really?
agent harnessSystemevery decision you can make without retraining the model — loop, context, tools, memory52The Harness
agent loopMechanismthink, act, look at the result, repeat until the job is done33From Answer to Action
AGIOpenQuestiona system that can do any cognitive job a person can — nobody agrees on the test43The Open Question
AI winterPhenomenona period when promises outran results and the money left17The First Winter
alignmentFieldmaking a system do what was meant, not merely what was written31Getting the Goal Right
any-to-anyPropertyone model that takes in text, image, audio or video and emits any of them54Omni-Models
APIProtocola fixed doorway one program uses to ask another for something34Hands and Function Calls
attentionMechanismlet each token weigh every other token, instead of only the last one24Attention Is All You Need
auxiliary-loss-free load balancingMechanismkeep experts evenly used by nudging a bias, instead of adding a penalty to the loss45DeepSeek
B
backpropagationAlgorithmassign blame for the error backwards through the layers, then nudge every weight10Rolling Downhill
backpropagation through timeAlgorithmunroll a loop over time, then run backprop through the whole tape21Memory in a Loop
barbellPhenomenonthe field splits toward a few huge closed models and many small open ones51The Open Frontier
batchingMechanismserving many requests together so the hardware is never idle38How Tokens Get Served
Bayes' ruleAlgorithmturn a prior and a likelihood into a posterior — updating belief with evidence08Belief, Updated
benchmarkBenchmarka fixed task suite and metric — the yardstick, and the thing people overfit to32How Do We Even Know?
biasQuantitya neuron's resting offset — how eager it is to fire with no input14The First Neuron
bias-variance tradeoffPropertyerror splits into being too simple and being too jumpy; lowering one raises the other11Why Memorizing Fails
byte-pair encodingAlgorithmbuild a vocabulary by repeatedly gluing together the most common pair of pieces25Breaking Language into Pieces
C
CANNSystemHuawei's software stack for Ascend chips — the counterpart to CUDA48The Compute Squeeze
capacityQuantityhow many different patterns a model is able to fit12No Free Lunch
chain-of-thoughtMechanismwriting the working out, which gives the model more room to be right41Thinking Before Answering
channelMechanismthe pipe a message travels through, which may corrupt it06Channels
channel capacityTheoremthe most information a noisy channel can carry — a hard ceiling06Channels
checkpointingMechanismsaving progress at safe points so a failure does not undo everything37Many Hands, One Job
ChinchillaResultfor a fixed budget, most big models of 2020 were too large and under-trained27More Is Different
chunkingMechanismcutting long documents into pieces small enough to retrieve and fit35Memory Outside the Weights
classifier-free guidanceMechanismsteer generation harder toward the prompt by contrasting with and without it29Noise into Images
code executionMechanismletting the model write and run real code instead of guessing the answer34Hands and Function Calls
combinatorial explosionFailureModeoptions multiply so fast that brute force stops being possible17The First Winter
commoditisationPhenomenoncapability that was rare becomes cheap, so value moves to what you build on top51The Open Frontier
compressionHeuristicintelligence is compression — predicting well and compressing well are one thing02Intelligence Is Compression
computabilityPropertywhether some step-by-step method can settle a question at all13Can Machines Think?
compute-optimalPropertythe best split of a fixed compute budget between model size and data27More Is Different
computer useMechanismdriving a real screen with mouse and keyboard, like a person would39A Safe Place to Act
constitutionDocumenta written set of principles a model is trained to critique itself against31Getting the Goal Right
contaminationFailureModetest questions that leaked into training, making the score a memory test32How Do We Even Know?
context managementMechanismdeciding what to keep, summarise or drop so the model sees the right things52The Harness
context windowPropertythe most text a model can hold in mind at once33From Answer to Action
control theoryFieldthe mathematics of steering a system toward a target without it wobbling18Feedback
ControlNetArchitecturea side branch that lets a sketch or pose steer an image model without retraining it53Image Models
convolutionMechanismslide one small filter across the input and reuse it everywhere — built-in translation symmetry20Seeing with Windows
convolutional neural networkArchitecturestacked convolutions that see images through learned local filters20Seeing with Windows
cross-entropyObjectivethe cost of coding data with the wrong model — the loss nearly everything trains on04Entropy
cyberneticsFieldthe 1940s study of control and communication in animals and machines alike18Feedback
D
denoisingMechanismlearning to remove a little noise, so many tiny steps turn noise into an image29Noise into Images
diffusion modelArchitecturelearn to undo noise, and you can make an image from nothing29Noise into Images
diffusion transformerArchitecturea transformer as the image denoiser, usually trained with flow matching53Image Models
distributional hypothesisHeuristicwords used in similar company end up meaning similar things22Words as Coordinates
dot productQuantitymultiply matching entries and add — a number for how aligned two vectors are09Everything Is a Vector
DPOAlgorithmlearn straight from preference pairs, skipping the separate reward model28Teaching Taste
driftPhenomenonthe world shifts, so a model that was right slowly stops being right40Watching the Loop
durable executionMechanisma long run that survives crashes and restarts where it left off37Many Hands, One Job
E
embedding spacePhenomenonthe geometry where similar things end up near each other09Everything Is a Vector
emergent abilityPhenomenona skill that barely shows, then appears suddenly as the model grows27More Is Different
environmentSystemeverything outside the agent that it can sense and change01What Is Intelligence, Really?
error-correcting codeAlgorithmadd redundancy so the receiver can undo the channel's damage06Channels
EU AI ActDocumentthe European Union law that sorts AI uses into risk tiers and sets duties for each49Licences and Rules
evaluationFielddeciding what a number means — validity, contamination, ablations32How Do We Even Know?
evaluation harnessSystemthe machinery that runs a suite of tests and scores the results the same way every time40Watching the Loop
expert systemSystema hand-built rulebook that answers like a specialist in a narrow domain16Rules All the Way Down
explorationMechanismdeliberately trying something new so you learn what you are missing23Learning from Reward
export controlsDocumentrules limiting chip sales across borders — the constraint that bent the field toward efficiency, then partly loosened48The Compute Squeeze
F
featureVariablea direction inside the model that stands for one human-readable idea30Opening the Black Box
feature mapQuantitythe grid of activations a filter produces as it slides across an image20Seeing with Windows
feedbackMechanismmeasure the gap to the goal and use it to correct the next step18Feedback
fine-tuningAlgorithmcontinue training a pretrained model on a smaller, targeted set28Teaching Taste
flow matchingMechanismlearn a straight path from noise to data instead of a slow curve53Image Models
FP8Formatan 8-bit number format for training and serving, roughly half the memory of 16-bit45DeepSeek
frame problemFailureModehard to say which facts an action leaves unchanged, without listing them all17The First Winter
full-duplexPropertylisten and speak at the same time, so talking feels like a conversation54Omni-Models
fully open modelPropertyopen weights plus open data, code and logs — reproducible end to end47The Other Open Lineages
function callingMechanismthe model asks for a function to be run, and reads the result back34Hands and Function Calls
G
generalizationPropertydoing well on data you have never seen — the only thing that counts02Intelligence Is Compression
GGUFFormata single self-contained file holding quantised weights and their metadata50Running One Yourself
GLMModelZhipu's model family, from GLM-130B through the GLM-4 line46The Chinese Labs
GPAIPropertygeneral-purpose AI — the EU category for large models usable for many tasks49Licences and Rules
GPUHardwarethousands of simple cores doing the same maths at once — the accident that trains everything15Why Silicon Got Good at This
gradient descentAlgorithmwalk downhill on the loss by following its slope10Rolling Downhill
GRPOAlgorithmscore a group of answers against each other, so no separate value network is needed45DeepSeek
guardrailMechanisma check that catches and blocks a bad output before a user sees it40Watching the Loop
H
halting problemResultno program can decide, for every program, whether it will ever stop13Can Machines Think?
harness engineeringFielddesigning the loop, tools, memory and permissions around a model — now a discipline of its own52The Harness
hidden layerMechanisma layer of neurons between input and output that builds its own features19Learning from Blame
Huawei AscendHardwarea non-Nvidia AI accelerator line, pushed hard by export controls48The Compute Squeeze
Huffman codingAlgorithmbuild the optimal prefix code by repeatedly merging the two least likely symbols05Codes
human-in-the-loopMechanisma person approves the risky step before the agent takes it39A Safe Place to Act
HunyuanModelTencent's open mixture-of-experts line, hundreds of billions of parameters46The Chinese Labs
I
inductive biasPropertythe assumptions a learner makes so it can guess about unseen cases12No Free Lunch
inferenceMechanismrunning a trained model to get an answer out38How Tokens Get Served
inference scalingPropertytrading more thinking at answer time for better answers41Thinking Before Answering
informationQuantityhow much uncertainty a message removes03The Bit
instruction tuningMechanismfine-tune on examples written as tasks, so following requests becomes the habit28Teaching Taste
intelligencePropertyreaching goals in an environment, across many different situations01What Is Intelligence, Really?
interconnectMechanismthe links between chips — often the real limit on training a large model48The Compute Squeeze
InternLMModelShanghai AI Laboratory's open model line, released with full technical reports46The Chinese Labs
interoperabilityPropertytools and agents from different makers working together without custom glue36A Port for Tools
J
JEPAArchitecturepredict in a learned abstract space, not pixel by pixel42Models of the World
K
KimiModelMoonshot's long-context line, with reasoning trained by large-scale RL46The Chinese Labs
KL divergenceQuantityextra bits wasted by believing q when the truth is p — never negative, never symmetric07Distance Between Beliefs
knowledge distillationMechanismtrain a small model to copy a big one's answers, gaining skill cheaply45DeepSeek
knowledge representationMechanismwriting what the system knows in a form a machine can reason over16Rules All the Way Down
KV cacheMechanismstore the keys and values you already computed, so each new token is cheap38How Tokens Get Served
L
language modelArchitecturea model that puts a probability on the next piece of text26The Big Read
latent diffusionArchitecturerun the diffusion steps in a small squeezed space, then decode back to pixels53Image Models
latent spacePropertya squeezed code where each point stands for a whole image, far smaller than pixels29Noise into Images
learning rateQuantityhow big a step to take downhill — too small is slow, too big overshoots10Rolling Downhill
least privilegeHeuristicgive every part the smallest power it needs to do the job39A Safe Place to Act
likelihoodQuantityhow probable the data you saw would be under a given hypothesis08Belief, Updated
LlamaModelthe family that made open weights mainstream — under a community licence with a user cap47The Other Open Lineages
local servingSystemrunning a model on your own machine, with no network round-trip50Running One Yourself
loop policyHeuristicthe rules for when to keep going, ask a human, or stop52The Harness
LoRAMechanismfreeze the model and train a tiny add-on — fine-tuning on one GPU50Running One Yourself
loss functionObjectivethe number training tries to make small — the formal statement of what you want10Rolling Downhill
LSTMArchitecturea gated loop that remembers over long stretches and eases the vanishing gradient21Memory in a Loop
M
Markov decision processMechanismstates, actions and rewards, where the next state depends only on now18Feedback
matrixQuantitya grid of numbers acting as one machine that turns vectors into vectors09Everything Is a Vector
MCP clientSystemthe agent side that connects to servers and offers their tools to the model36A Port for Tools
MCP serverSystema small program that exposes tools and data through the shared protocol36A Port for Tools
mechanistic interpretabilityFieldtaking the model apart to name the parts and what each one does30Opening the Black Box
memory bandwidthQuantityhow fast numbers move from memory — the limit that decoding usually hits first15Why Silicon Got Good at This
memory wallPhenomenoncompute got fast but memory did not, so moving numbers is now the bottleneck48The Compute Squeeze
MiniMaxModela Chinese lab building long-context models on lightning attention and mixture of experts46The Chinese Labs
MistralModela European lab whose small open models pushed quality per parameter47The Other Open Lineages
mixture of expertsMechanismroute each token to a few small experts — huge capacity, small active cost45DeepSeek
modality encoderArchitecturethe part that turns one kind of input — sound, image, video — into shared tokens54Omni-Models
model cardDocumenta short factual page describing what a model is, how it was made and its limits49Licences and Rules
Model Context ProtocolProtocola standard way to plug tools into a model, so every model need not learn every tool36A Port for Tools
model familyPropertyone line of models at several sizes, sharing a recipe and a name46The Chinese Labs
model-based RLMechanismlearn a model of the world, then plan inside it instead of acting blindly42Models of the World
multi-agentSystemseveral agents splitting a job, talking to each other as they go37Many Hands, One Job
multi-head latent attentionMechanismcompress the per-token memory attention must keep, so long contexts cost far less45DeepSeek
multi-token predictionMechanismtrain the model to guess several next tokens at once, not just one45DeepSeek
multilayer perceptronArchitecturestacked neuron layers that can bend a boundary no single perceptron could19Learning from Blame
mutual informationQuantityhow much knowing one thing tells you about another07Distance Between Beliefs
N
neuronMechanismadd the inputs, then pass the total through a simple curve14The First Neuron
next-token predictionTaskthe one training game behind a language model: guess the next token26The Big Read
no free lunchTheoremaveraged over all problems, no learner beats any other — so assumptions are everything12No Free Lunch
noisePhenomenonrandom damage the channel adds to the message06Channels
O
observabilityFieldbeing able to see what a running system is doing and why it failed40Watching the Loop
Occam's razorHeuristicprefer the simplest explanation that fits — a guess, not a theorem12No Free Lunch
OLMoModela fully open model line that releases data, code and training logs too47The Other Open Lineages
omni-modelArchitectureone model for text, audio, images and video, instead of a relay of specialists54Omni-Models
open image modelModelan image generator you can download and run yourself53Image Models
open weightsPropertythe numbers released, letting anyone run and fine-tune — but not rebuild44What 'Open' Means
open-source AIFieldbuilding AI in the open, where anyone can read, run and change it44What 'Open' Means
open-weight baselineBenchmarkthe best open model, used as the yardstick everyone else compares against47The Other Open Lineages
open-weight licenceDocumentthe terms attached to downloadable weights — open weights are not always open source49Licences and Rules
open-weight riskOpenQuestiondoes releasing weights let the dangerous uses scale faster than the defensive ones?51The Open Frontier
openness ladderPropertya scale from weights-only to fully open data, code and logs44What 'Open' Means
orchestrationSystemthe layer that decides what runs, in what order, and what happens when it fails37Many Hands, One Job
overfittingFailureModememorising the training set instead of learning the pattern11Why Memorizing Fails
P
parallelismPropertydoing many small sums at once instead of one after another15Why Silicon Got Good at This
perceptronMechanismthe first machine that learned weights to separate two classes14The First Neuron
permissionMechanismthe grant that lets a program read this file or call that service — and nothing else39A Safe Place to Act
permissive licenceDocumenta licence that lets you use, change and sell with almost no conditions44What 'Open' Means
perplexityQuantityhow many equally likely choices the model felt it was facing — cross-entropy, exponentiated07Distance Between Beliefs
planningMechanismchoosing a whole sequence of steps in advance of doing them33From Answer to Action
policyMechanismthe rule that says which action to take in each situation23Learning from Reward
poolingMechanismshrink a feature map by keeping the strongest value in each patch20Seeing with Windows
positional encodingMechanisma pattern added so the model knows where each token sat in the order24Attention Is All You Need
posteriorDistributionwhat you believe after seeing the evidence08Belief, Updated
predictionMechanismsaying what comes next before you are told02Intelligence Is Compression
prefix codeMechanisma code where no codeword is the start of another, so it is readable as it arrives05Codes
pretrainingAlgorithmlearn from raw text at scale by predicting what comes next26The Big Read
priorDistributionwhat you believed before seeing the evidence08Belief, Updated
probabilityFieldthe mathematics of belief under uncertainty08Belief, Updated
Q
quantizationMechanismstore weights in fewer bits, so a model fits on hardware you own38How Tokens Get Served
quantized weightsFormatweights stored at 8, 4 or fewer bits so a big model fits in small memory50Running One Yourself
query-key-valueMechanismeach token asks a question, every token advertises what it holds, and answers are blended by fit24Attention Is All You Need
QwenModelAlibaba's open model line, released across many sizes and modalities46The Chinese Labs
R
ReActAlgorithminterleave a line of reasoning with a tool call, so thought follows evidence33From Answer to Action
reasoning modelArchitecturea model trained to spend more thinking before it answers41Thinking Before Answering
recurrent neural networkArchitecturea network that carries a hidden state forward, one step at a time21Memory in a Loop
red teamingMechanismattacking the system on purpose to find how it breaks before others do32How Do We Even Know?
regularizationMechanismany penalty that keeps a model simple so it stops memorising11Why Memorizing Fails
reinforcement learningFieldlearning what to do by trying, from delayed reward23Learning from Reward
retrieval-augmented generationAlgorithmlook facts up and put them in the context, instead of storing them in the weights35Memory Outside the Weights
rewardObjectivea single number saying how well things just went01What Is Intelligence, Really?
reward hackingFailureModegetting the reward without doing the thing — the optimiser finds the loophole31Getting the Goal Right
reward modelModela model trained on human preferences that scores answers so training can use it28Teaching Taste
RLHFAlgorithmtilt a pretrained model toward human preference — how raw prediction becomes behaviour28Teaching Taste
S
sandboxMechanisman isolated place for an agent to act, so a wrong step is survivable39A Safe Place to Act
scaffoldingSystemevery decision you can make without retraining the model — loop, context, tools, memory52The Harness
scaling debateOpenQuestionis more scale enough for general capability, or is something still missing?43The Open Question
scaling lawResultloss falls as a smooth power law in compute, data and parameters — so you can plan ahead27More Is Different
searchAlgorithmtry possibilities in a smart order instead of all of them16Rules All the Way Down
self-attentionMechanismlet each token weigh every other token, instead of only the last one24Attention Is All You Need
self-supervised learningMechanismhide part of the data and let the model teach itself to recover it26The Big Read
self-supervised videoMechanismlearn from raw video by predicting what happens next in it42Models of the World
semantic searchMechanismfinding things by meaning rather than by matching words35Memory Outside the Weights
sequence modelArchitectureany model built to read or write things in order21Memory in a Loop
source codingTheoremyou cannot compress below the entropy — and you can get arbitrarily close05Codes
sparse autoencoderArchitecturea tool that pulls many tidy, mostly-silent features out of a crowded layer30Opening the Black Box
specificationDocumentwriting down what we actually want the system to do31Getting the Goal Right
speculative decodingAlgorithma small fast model drafts, a big model checks — same output, fewer big steps38How Tokens Get Served
speech tokenQuantitya chunk of sound turned into something the same machinery can treat like a word54Omni-Models
superintelligenceOpenQuestiona system smarter than us in every way — a hypothesis, not an observation43The Open Question
superpositionPhenomenona model packs more ideas than it has directions, by sharing them at angles30Opening the Black Box
symbolic AIFieldintelligence as rules over written symbols, added by hand16Rules All the Way Down
T
tensor coreHardwarea circuit that does a whole small matrix multiply in one instruction15Why Silicon Got Good at This
test setDatasetexamples held back, used only to check whether learning generalised11Why Memorizing Fails
test-time computeQuantitythe effort spent while answering, not while training41Thinking Before Answering
tokenQuantitya chunk of text — often a word-piece — that the model reads as one unit25Breaking Language into Pieces
tokenizerMechanismthe rule that cuts text into tokens — it shapes everything downstream25Breaking Language into Pieces
tool schemaFormata written description of a function — its name, inputs and outputs — a model can be handed34Hands and Function Calls
tracingMechanismrecording every step, input and output so a run can be replayed40Watching the Loop
training setDatasetthe examples the model is allowed to learn from11Why Memorizing Fails
transformerArchitectureattention and feed-forward blocks stacked — the architecture that swallowed the field24Attention Is All You Need
Turing machineMechanisman imagined tape-reading machine that pins down what "computable" means13Can Machines Think?
Turing testThoughtExperimentcan you tell the machine from a person in conversation? — a test of behaviour, not mind13Can Machines Think?
U
uncertaintyPropertyhow many answers are still possible03The Bit
underfittingFailureModethe model is too simple to capture the pattern, so it is wrong everywhere11Why Memorizing Fails
V
value functionEstimatora guess of how much reward is still to come from here23Learning from Reward
vanishing gradientFailureModein a deep stack the learning signal fades to nothing before it reaches the bottom19Learning from Blame
vectorQuantitya list of numbers treated as one point or one arrow09Everything Is a Vector
vector arithmeticPhenomenonking − man + woman ≈ queen: directions in embedding space carry meaning22Words as Coordinates
vector databaseSystema store that finds the nearest vectors fast, even among billions35Memory Outside the Weights
vLLMSystema serving engine whose paged memory made long-context batching practical38How Tokens Get Served
vocabularyDatasetthe fixed set of tokens a model can read and write25Breaking Language into Pieces
W
weightQuantityhow much a neuron cares about one input — the number training changes14The First Neuron
word embeddingMechanismwords as coordinates, so that similar meanings sit near each other22Words as Coordinates
word2vecAlgorithmlearn a vector per word by predicting its neighbours, cheaply and at scale22Words as Coordinates
workflowMechanisma fixed sequence of steps, with the model doing only the fuzzy parts37Many Hands, One Job
world modelArchitecturean internal simulator of how a situation evolves — imagining without acting42Models of the World
No term matches that.