AI Foundationspredict · compress · act
act IV

The Machines

Neural networks are mostly one operation, repeated. The GPU was built for exactly that operation.

15

Why Silicon Got Good at This

before this →The First Neuron
2007CUDA — NVIDIA opens general-purpose programming on graphics chips — the accident that trains everything since.
2012AlexNet — the thaw, and the boom that never stopped

Deep learning’s core computation is the dot product, repeated billions of times. The CPU is a handful of very clever workers. The GPU is thousands of simple ones. For this workload, thousands of simple beats a few clever.

CPUcoreGPUthousands of threads
CPU: few fast cores. GPU: thousands of small cores plus wide memory.

Two bottlenecks, not one

Raw arithmetic is only half the story. A tensor core can multiply small matrices in a single clock tick, which is why modern chips advertise enormous “FLOPs.” But the real limit is often memory bandwidth: how fast numbers can be fed to those cores. A GPU that spends most of its time waiting on memory is idle. This is why so much engineering goes into caches, fusion, and keeping data close to compute.

Why this mattered so much

Nothing about deep learning changed in 2012 except scale. In 2009 a GPU could train a network roughly a hundred times faster than a CPU. That is not a speed-up; it is a different universe of what is reachable. Every large model since is downstream of this one fact: the hardware for matrix math got cheap and absurdly fast.

Parallelism is the whole trick

same math, different shape

Training a network is embarrassingly parallel: every example in a batch can be processed at once, and every neuron in a layer can fire at once. The GPU exists to do many independent things simultaneously — which is exactly what a matrix multiply is.

the patternAI progress has three fuels: algorithms, data, and compute. From here on, compute is the one that sets the ceiling.
MATRIX MATHthe same operation, billions of times
THROUGHPUTmany cores, many examples at once
SCALEwhat was impossible becomes routine
introduces →parallelismGPUtensor corememory bandwidth
← previousThe First Neuronnext →Rules All the Way Down