AI Foundationspredict · compress · act

The cross-index

The same chapters, sorted by the concept each introduces. This is how the document builds on itself.

Prerequisite chain

01What Is Intelligence, Really?root
02Intelligence Is CompressionWhat Is Intelligence, Really?
03The Bit: A Yes or a NoIntelligence Is Compression
04Entropy: The Price of SurpriseThe Bit: A Yes or a No
05Codes: Saying More with LessEntropy: The Price of Surprise
06Channels: Pushing Through NoiseCodes: Saying More with Less
07Distance Between BeliefsChannels: Pushing Through Noise
08Belief, UpdatedDistance Between Beliefs
09Everything Is a VectorBelief, Updated
10Rolling DownhillEverything Is a Vector
11Why Memorizing FailsRolling Downhill
12No Free LunchWhy Memorizing Fails
13Can Machines Think?No Free Lunch
14The First NeuronCan Machines Think?
15Why Silicon Got Good at ThisThe First Neuron
16Rules All the Way DownWhy Silicon Got Good at This
17The First WinterRules All the Way Down
18Feedback: The Other TreeThe First Winter
19Learning from BlameFeedback: The Other Tree
20Seeing with WindowsLearning from Blame
21Memory in a LoopSeeing with Windows
22Words as CoordinatesMemory in a Loop
23Learning from RewardWords as Coordinates
24Attention Is All You NeedLearning from Reward
25Breaking Language into PiecesAttention Is All You Need
26The Big ReadBreaking Language into Pieces
27More Is DifferentThe Big Read
28Teaching TasteMore Is Different
29Noise into ImagesTeaching Taste
30Opening the Black BoxNoise into Images
31Getting the Goal RightOpening the Black Box
32How Do We Even Know?Getting the Goal Right
33From Answer to ActionHow Do We Even Know?
34Hands and Function CallsFrom Answer to Action
35Memory Outside the WeightsHands and Function Calls
36A Port for ToolsMemory Outside the Weights
37Many Hands, One JobA Port for Tools
38How Tokens Get ServedMany Hands, One Job
39A Safe Place to ActHow Tokens Get Served
40Watching the LoopA Safe Place to Act
41Thinking Before AnsweringWatching the Loop
42Models of the WorldThinking Before Answering
43The Open QuestionModels of the World
44What 'Open' MeansThe Open Question
45DeepSeekWhat 'Open' Means
46The Chinese LabsDeepSeek
47The Other Open LineagesWhat 'Open' Means
48The Compute SqueezeThe Chinese Labs
49Licences and RulesThe Compute Squeeze
50Running One YourselfLicences and Rules
51The Open FrontierRunning One Yourself
52The HarnessWatching the Loop
53Image ModelsNoise into Images
54Omni-ModelsImage Models

Concepts → chapters

▪intelligenceWhat Is Intelligence, Really?
▪agentWhat Is Intelligence, Really?
▪environmentWhat Is Intelligence, Really?
▪rewardWhat Is Intelligence, Really?
▪compressionIntelligence Is Compression
▪predictionIntelligence Is Compression
▪generalizationIntelligence Is Compression
▪bitThe Bit: A Yes or a No
▪uncertaintyThe Bit: A Yes or a No
▪informationThe Bit: A Yes or a No
▪entropyEntropy: The Price of Surprise
▪surpriseEntropy: The Price of Surprise
▪cross-entropyEntropy: The Price of Surprise
▪source codingCodes: Saying More with Less
▪prefix codeCodes: Saying More with Less
▪Huffman codingCodes: Saying More with Less
▪rateCodes: Saying More with Less
▪channelChannels: Pushing Through Noise
▪noiseChannels: Pushing Through Noise
▪channel capacityChannels: Pushing Through Noise
▪error-correcting codeChannels: Pushing Through Noise
▪KL divergenceDistance Between Beliefs
▪mutual informationDistance Between Beliefs
▪perplexityDistance Between Beliefs
▪probabilityBelief, Updated
▪Bayes' ruleBelief, Updated
▪priorBelief, Updated
▪posteriorBelief, Updated
▪likelihoodBelief, Updated
▪vectorEverything Is a Vector
▪dot productEverything Is a Vector
▪matrixEverything Is a Vector
▪embedding spaceEverything Is a Vector
▪loss functionRolling Downhill
▪gradient descentRolling Downhill
▪backpropagationRolling Downhill
▪learning rateRolling Downhill
▪overfittingWhy Memorizing Fails
▪underfittingWhy Memorizing Fails
▪regularizationWhy Memorizing Fails
▪training setWhy Memorizing Fails
▪test setWhy Memorizing Fails
▪bias-variance tradeoffWhy Memorizing Fails
▪inductive biasNo Free Lunch
▪no free lunchNo Free Lunch
▪Occam's razorNo Free Lunch
▪capacityNo Free Lunch
▪Turing machineCan Machines Think?
▪computabilityCan Machines Think?
▪Turing testCan Machines Think?
▪halting problemCan Machines Think?
▪neuronThe First Neuron
▪perceptronThe First Neuron
▪activation functionThe First Neuron
▪weightThe First Neuron
▪biasThe First Neuron
▪parallelismWhy Silicon Got Good at This
▪GPUWhy Silicon Got Good at This
▪tensor coreWhy Silicon Got Good at This
▪memory bandwidthWhy Silicon Got Good at This
▪symbolic AIRules All the Way Down
▪knowledge representationRules All the Way Down
▪expert systemRules All the Way Down
▪searchRules All the Way Down
▪AI winterThe First Winter
▪combinatorial explosionThe First Winter
▪frame problemThe First Winter
▪feedbackFeedback: The Other Tree
▪control theoryFeedback: The Other Tree
▪Markov decision processFeedback: The Other Tree
▪cyberneticsFeedback: The Other Tree
▪hidden layerLearning from Blame
▪multilayer perceptronLearning from Blame
▪vanishing gradientLearning from Blame
▪convolutionSeeing with Windows
▪convolutional neural networkSeeing with Windows
▪feature mapSeeing with Windows
▪poolingSeeing with Windows
▪recurrent neural networkMemory in a Loop
▪LSTMMemory in a Loop
▪sequence modelMemory in a Loop
▪backpropagation through timeMemory in a Loop
▪word embeddingWords as Coordinates
▪word2vecWords as Coordinates
▪distributional hypothesisWords as Coordinates
▪vector arithmeticWords as Coordinates
▪reinforcement learningLearning from Reward
▪policyLearning from Reward
▪value functionLearning from Reward
▪explorationLearning from Reward
▪attentionAttention Is All You Need
▪transformerAttention Is All You Need
▪self-attentionAttention Is All You Need
▪query-key-valueAttention Is All You Need
▪positional encodingAttention Is All You Need
▪tokenBreaking Language into Pieces
▪tokenizerBreaking Language into Pieces
▪byte-pair encodingBreaking Language into Pieces
▪vocabularyBreaking Language into Pieces
▪pretrainingThe Big Read
▪self-supervised learningThe Big Read
▪language modelThe Big Read
▪next-token predictionThe Big Read
▪scaling lawMore Is Different
▪compute-optimalMore Is Different
▪emergent abilityMore Is Different
▪ChinchillaMore Is Different
▪fine-tuningTeaching Taste
▪instruction tuningTeaching Taste
▪RLHFTeaching Taste
▪reward modelTeaching Taste
▪DPOTeaching Taste
▪diffusion modelNoise into Images
▪denoisingNoise into Images
▪latent spaceNoise into Images
▪classifier-free guidanceNoise into Images
▪mechanistic interpretabilityOpening the Black Box
▪featureOpening the Black Box
▪superpositionOpening the Black Box
▪sparse autoencoderOpening the Black Box
▪alignmentGetting the Goal Right
▪specificationGetting the Goal Right
▪reward hackingGetting the Goal Right
▪constitutionGetting the Goal Right
▪benchmarkHow Do We Even Know?
▪evaluationHow Do We Even Know?
▪contaminationHow Do We Even Know?
▪red teamingHow Do We Even Know?
▪agent loopFrom Answer to Action
▪ReActFrom Answer to Action
▪planningFrom Answer to Action
▪context windowFrom Answer to Action
▪function callingHands and Function Calls
▪tool schemaHands and Function Calls
▪code executionHands and Function Calls
▪APIHands and Function Calls
▪retrieval-augmented generationMemory Outside the Weights
▪vector databaseMemory Outside the Weights
▪chunkingMemory Outside the Weights
▪semantic searchMemory Outside the Weights
▪Model Context ProtocolA Port for Tools
▪MCP serverA Port for Tools
▪MCP clientA Port for Tools
▪interoperabilityA Port for Tools
▪orchestrationMany Hands, One Job
▪workflowMany Hands, One Job
▪multi-agentMany Hands, One Job
▪durable executionMany Hands, One Job
▪checkpointingMany Hands, One Job
▪inferenceHow Tokens Get Served
▪KV cacheHow Tokens Get Served
▪batchingHow Tokens Get Served
▪quantizationHow Tokens Get Served
▪vLLMHow Tokens Get Served
▪speculative decodingHow Tokens Get Served
▪sandboxA Safe Place to Act
▪permissionA Safe Place to Act
▪least privilegeA Safe Place to Act
▪computer useA Safe Place to Act
▪human-in-the-loopA Safe Place to Act
▪tracingWatching the Loop
▪observabilityWatching the Loop
▪evaluation harnessWatching the Loop
▪driftWatching the Loop
▪guardrailWatching the Loop
▪chain-of-thoughtThinking Before Answering
▪reasoning modelThinking Before Answering
▪test-time computeThinking Before Answering
▪inference scalingThinking Before Answering
▪world modelModels of the World
▪model-based RLModels of the World
▪self-supervised videoModels of the World
▪JEPAModels of the World
▪AGIThe Open Question
▪superintelligenceThe Open Question
▪scaling debateThe Open Question
▪open weightsWhat 'Open' Means
▪open-source AIWhat 'Open' Means
▪openness ladderWhat 'Open' Means
▪permissive licenceWhat 'Open' Means
▪mixture of expertsDeepSeek
▪multi-head latent attentionDeepSeek
▪auxiliary-loss-free load balancingDeepSeek
▪multi-token predictionDeepSeek
▪GRPODeepSeek
▪FP8DeepSeek
▪knowledge distillationDeepSeek
▪model familyThe Chinese Labs
▪QwenThe Chinese Labs
▪GLMThe Chinese Labs
▪KimiThe Chinese Labs
▪MiniMaxThe Chinese Labs
▪InternLMThe Chinese Labs
▪HunyuanThe Chinese Labs
▪LlamaThe Other Open Lineages
▪MistralThe Other Open Lineages
▪OLMoThe Other Open Lineages
▪fully open modelThe Other Open Lineages
▪open-weight baselineThe Other Open Lineages
▪export controlsThe Compute Squeeze
▪Huawei AscendThe Compute Squeeze
▪CANNThe Compute Squeeze
▪interconnectThe Compute Squeeze
▪memory wallThe Compute Squeeze
▪open-weight licenceLicences and Rules
▪acceptable use policyLicences and Rules
▪model cardLicences and Rules
▪EU AI ActLicences and Rules
▪GPAILicences and Rules
▪GGUFRunning One Yourself
▪LoRARunning One Yourself
▪local servingRunning One Yourself
▪quantized weightsRunning One Yourself
▪commoditisationThe Open Frontier
▪open-weight riskThe Open Frontier
▪barbellThe Open Frontier
▪agent harnessThe Harness
▪harness engineeringThe Harness
▪context managementThe Harness
▪loop policyThe Harness
▪scaffoldingThe Harness
▪latent diffusionImage Models
▪diffusion transformerImage Models
▪flow matchingImage Models
▪ControlNetImage Models
▪open image modelImage Models
▪omni-modelOmni-Models
▪any-to-anyOmni-Models
▪full-duplexOmni-Models
▪speech tokenOmni-Models
▪modality encoderOmni-Models