AI Foundations
predict · compress · act
contents
cross-index
history
lineages
glossary
The cross-index
The same chapters, sorted by the concept each introduces. This is how the document builds on itself.
Prerequisite chain
01
What Is Intelligence, Really?
root
02
Intelligence Is Compression
What Is Intelligence, Really?
03
The Bit: A Yes or a No
Intelligence Is Compression
04
Entropy: The Price of Surprise
The Bit: A Yes or a No
05
Codes: Saying More with Less
Entropy: The Price of Surprise
06
Channels: Pushing Through Noise
Codes: Saying More with Less
07
Distance Between Beliefs
Channels: Pushing Through Noise
08
Belief, Updated
Distance Between Beliefs
09
Everything Is a Vector
Belief, Updated
10
Rolling Downhill
Everything Is a Vector
11
Why Memorizing Fails
Rolling Downhill
12
No Free Lunch
Why Memorizing Fails
13
Can Machines Think?
No Free Lunch
14
The First Neuron
Can Machines Think?
15
Why Silicon Got Good at This
The First Neuron
16
Rules All the Way Down
Why Silicon Got Good at This
17
The First Winter
Rules All the Way Down
18
Feedback: The Other Tree
The First Winter
19
Learning from Blame
Feedback: The Other Tree
20
Seeing with Windows
Learning from Blame
21
Memory in a Loop
Seeing with Windows
22
Words as Coordinates
Memory in a Loop
23
Learning from Reward
Words as Coordinates
24
Attention Is All You Need
Learning from Reward
25
Breaking Language into Pieces
Attention Is All You Need
26
The Big Read
Breaking Language into Pieces
27
More Is Different
The Big Read
28
Teaching Taste
More Is Different
29
Noise into Images
Teaching Taste
30
Opening the Black Box
Noise into Images
31
Getting the Goal Right
Opening the Black Box
32
How Do We Even Know?
Getting the Goal Right
33
From Answer to Action
How Do We Even Know?
34
Hands and Function Calls
From Answer to Action
35
Memory Outside the Weights
Hands and Function Calls
36
A Port for Tools
Memory Outside the Weights
37
Many Hands, One Job
A Port for Tools
38
How Tokens Get Served
Many Hands, One Job
39
A Safe Place to Act
How Tokens Get Served
40
Watching the Loop
A Safe Place to Act
41
Thinking Before Answering
Watching the Loop
42
Models of the World
Thinking Before Answering
43
The Open Question
Models of the World
44
What 'Open' Means
The Open Question
45
DeepSeek
What 'Open' Means
46
The Chinese Labs
DeepSeek
47
The Other Open Lineages
What 'Open' Means
48
The Compute Squeeze
The Chinese Labs
49
Licences and Rules
The Compute Squeeze
50
Running One Yourself
Licences and Rules
51
The Open Frontier
Running One Yourself
52
The Harness
Watching the Loop
53
Image Models
Noise into Images
54
Omni-Models
Image Models
Concepts → chapters
▪
intelligence
What Is Intelligence, Really?
▪
agent
What Is Intelligence, Really?
▪
environment
What Is Intelligence, Really?
▪
reward
What Is Intelligence, Really?
▪
compression
Intelligence Is Compression
▪
prediction
Intelligence Is Compression
▪
generalization
Intelligence Is Compression
▪
bit
The Bit: A Yes or a No
▪
uncertainty
The Bit: A Yes or a No
▪
information
The Bit: A Yes or a No
▪
entropy
Entropy: The Price of Surprise
▪
surprise
Entropy: The Price of Surprise
▪
cross-entropy
Entropy: The Price of Surprise
▪
source coding
Codes: Saying More with Less
▪
prefix code
Codes: Saying More with Less
▪
Huffman coding
Codes: Saying More with Less
▪
rate
Codes: Saying More with Less
▪
channel
Channels: Pushing Through Noise
▪
noise
Channels: Pushing Through Noise
▪
channel capacity
Channels: Pushing Through Noise
▪
error-correcting code
Channels: Pushing Through Noise
▪
KL divergence
Distance Between Beliefs
▪
mutual information
Distance Between Beliefs
▪
perplexity
Distance Between Beliefs
▪
probability
Belief, Updated
▪
Bayes' rule
Belief, Updated
▪
prior
Belief, Updated
▪
posterior
Belief, Updated
▪
likelihood
Belief, Updated
▪
vector
Everything Is a Vector
▪
dot product
Everything Is a Vector
▪
matrix
Everything Is a Vector
▪
embedding space
Everything Is a Vector
▪
loss function
Rolling Downhill
▪
gradient descent
Rolling Downhill
▪
backpropagation
Rolling Downhill
▪
learning rate
Rolling Downhill
▪
overfitting
Why Memorizing Fails
▪
underfitting
Why Memorizing Fails
▪
regularization
Why Memorizing Fails
▪
training set
Why Memorizing Fails
▪
test set
Why Memorizing Fails
▪
bias-variance tradeoff
Why Memorizing Fails
▪
inductive bias
No Free Lunch
▪
no free lunch
No Free Lunch
▪
Occam's razor
No Free Lunch
▪
capacity
No Free Lunch
▪
Turing machine
Can Machines Think?
▪
computability
Can Machines Think?
▪
Turing test
Can Machines Think?
▪
halting problem
Can Machines Think?
▪
neuron
The First Neuron
▪
perceptron
The First Neuron
▪
activation function
The First Neuron
▪
weight
The First Neuron
▪
bias
The First Neuron
▪
parallelism
Why Silicon Got Good at This
▪
GPU
Why Silicon Got Good at This
▪
tensor core
Why Silicon Got Good at This
▪
memory bandwidth
Why Silicon Got Good at This
▪
symbolic AI
Rules All the Way Down
▪
knowledge representation
Rules All the Way Down
▪
expert system
Rules All the Way Down
▪
search
Rules All the Way Down
▪
AI winter
The First Winter
▪
combinatorial explosion
The First Winter
▪
frame problem
The First Winter
▪
feedback
Feedback: The Other Tree
▪
control theory
Feedback: The Other Tree
▪
Markov decision process
Feedback: The Other Tree
▪
cybernetics
Feedback: The Other Tree
▪
hidden layer
Learning from Blame
▪
multilayer perceptron
Learning from Blame
▪
vanishing gradient
Learning from Blame
▪
convolution
Seeing with Windows
▪
convolutional neural network
Seeing with Windows
▪
feature map
Seeing with Windows
▪
pooling
Seeing with Windows
▪
recurrent neural network
Memory in a Loop
▪
LSTM
Memory in a Loop
▪
sequence model
Memory in a Loop
▪
backpropagation through time
Memory in a Loop
▪
word embedding
Words as Coordinates
▪
word2vec
Words as Coordinates
▪
distributional hypothesis
Words as Coordinates
▪
vector arithmetic
Words as Coordinates
▪
reinforcement learning
Learning from Reward
▪
policy
Learning from Reward
▪
value function
Learning from Reward
▪
exploration
Learning from Reward
▪
attention
Attention Is All You Need
▪
transformer
Attention Is All You Need
▪
self-attention
Attention Is All You Need
▪
query-key-value
Attention Is All You Need
▪
positional encoding
Attention Is All You Need
▪
token
Breaking Language into Pieces
▪
tokenizer
Breaking Language into Pieces
▪
byte-pair encoding
Breaking Language into Pieces
▪
vocabulary
Breaking Language into Pieces
▪
pretraining
The Big Read
▪
self-supervised learning
The Big Read
▪
language model
The Big Read
▪
next-token prediction
The Big Read
▪
scaling law
More Is Different
▪
compute-optimal
More Is Different
▪
emergent ability
More Is Different
▪
Chinchilla
More Is Different
▪
fine-tuning
Teaching Taste
▪
instruction tuning
Teaching Taste
▪
RLHF
Teaching Taste
▪
reward model
Teaching Taste
▪
DPO
Teaching Taste
▪
diffusion model
Noise into Images
▪
denoising
Noise into Images
▪
latent space
Noise into Images
▪
classifier-free guidance
Noise into Images
▪
mechanistic interpretability
Opening the Black Box
▪
feature
Opening the Black Box
▪
superposition
Opening the Black Box
▪
sparse autoencoder
Opening the Black Box
▪
alignment
Getting the Goal Right
▪
specification
Getting the Goal Right
▪
reward hacking
Getting the Goal Right
▪
constitution
Getting the Goal Right
▪
benchmark
How Do We Even Know?
▪
evaluation
How Do We Even Know?
▪
contamination
How Do We Even Know?
▪
red teaming
How Do We Even Know?
▪
agent loop
From Answer to Action
▪
ReAct
From Answer to Action
▪
planning
From Answer to Action
▪
context window
From Answer to Action
▪
function calling
Hands and Function Calls
▪
tool schema
Hands and Function Calls
▪
code execution
Hands and Function Calls
▪
API
Hands and Function Calls
▪
retrieval-augmented generation
Memory Outside the Weights
▪
vector database
Memory Outside the Weights
▪
chunking
Memory Outside the Weights
▪
semantic search
Memory Outside the Weights
▪
Model Context Protocol
A Port for Tools
▪
MCP server
A Port for Tools
▪
MCP client
A Port for Tools
▪
interoperability
A Port for Tools
▪
orchestration
Many Hands, One Job
▪
workflow
Many Hands, One Job
▪
multi-agent
Many Hands, One Job
▪
durable execution
Many Hands, One Job
▪
checkpointing
Many Hands, One Job
▪
inference
How Tokens Get Served
▪
KV cache
How Tokens Get Served
▪
batching
How Tokens Get Served
▪
quantization
How Tokens Get Served
▪
vLLM
How Tokens Get Served
▪
speculative decoding
How Tokens Get Served
▪
sandbox
A Safe Place to Act
▪
permission
A Safe Place to Act
▪
least privilege
A Safe Place to Act
▪
computer use
A Safe Place to Act
▪
human-in-the-loop
A Safe Place to Act
▪
tracing
Watching the Loop
▪
observability
Watching the Loop
▪
evaluation harness
Watching the Loop
▪
drift
Watching the Loop
▪
guardrail
Watching the Loop
▪
chain-of-thought
Thinking Before Answering
▪
reasoning model
Thinking Before Answering
▪
test-time compute
Thinking Before Answering
▪
inference scaling
Thinking Before Answering
▪
world model
Models of the World
▪
model-based RL
Models of the World
▪
self-supervised video
Models of the World
▪
JEPA
Models of the World
▪
AGI
The Open Question
▪
superintelligence
The Open Question
▪
scaling debate
The Open Question
▪
open weights
What 'Open' Means
▪
open-source AI
What 'Open' Means
▪
openness ladder
What 'Open' Means
▪
permissive licence
What 'Open' Means
▪
mixture of experts
DeepSeek
▪
multi-head latent attention
DeepSeek
▪
auxiliary-loss-free load balancing
DeepSeek
▪
multi-token prediction
DeepSeek
▪
GRPO
DeepSeek
▪
FP8
DeepSeek
▪
knowledge distillation
DeepSeek
▪
model family
The Chinese Labs
▪
Qwen
The Chinese Labs
▪
GLM
The Chinese Labs
▪
Kimi
The Chinese Labs
▪
MiniMax
The Chinese Labs
▪
InternLM
The Chinese Labs
▪
Hunyuan
The Chinese Labs
▪
Llama
The Other Open Lineages
▪
Mistral
The Other Open Lineages
▪
OLMo
The Other Open Lineages
▪
fully open model
The Other Open Lineages
▪
open-weight baseline
The Other Open Lineages
▪
export controls
The Compute Squeeze
▪
Huawei Ascend
The Compute Squeeze
▪
CANN
The Compute Squeeze
▪
interconnect
The Compute Squeeze
▪
memory wall
The Compute Squeeze
▪
open-weight licence
Licences and Rules
▪
acceptable use policy
Licences and Rules
▪
model card
Licences and Rules
▪
EU AI Act
Licences and Rules
▪
GPAI
Licences and Rules
▪
GGUF
Running One Yourself
▪
LoRA
Running One Yourself
▪
local serving
Running One Yourself
▪
quantized weights
Running One Yourself
▪
commoditisation
The Open Frontier
▪
open-weight risk
The Open Frontier
▪
barbell
The Open Frontier
▪
agent harness
The Harness
▪
harness engineering
The Harness
▪
context management
The Harness
▪
loop policy
The Harness
▪
scaffolding
The Harness
▪
latent diffusion
Image Models
▪
diffusion transformer
Image Models
▪
flow matching
Image Models
▪
ControlNet
Image Models
▪
open image model
Image Models
▪
omni-model
Omni-Models
▪
any-to-any
Omni-Models
▪
full-duplex
Omni-Models
▪
speech token
Omni-Models
▪
modality encoder
Omni-Models