Why Silicon Got Good at This
Deep learning’s core computation is the dot product, repeated billions of times. The CPU is a handful of very clever workers. The GPU is thousands of simple ones. For this workload, thousands of simple beats a few clever.
Two bottlenecks, not one
Raw arithmetic is only half the story. A tensor core can multiply small matrices in a single clock tick, which is why modern chips advertise enormous “FLOPs.” But the real limit is often memory bandwidth: how fast numbers can be fed to those cores. A GPU that spends most of its time waiting on memory is idle. This is why so much engineering goes into caches, fusion, and keeping data close to compute.
Why this mattered so much
Nothing about deep learning changed in 2012 except scale. In 2009 a GPU could train a network roughly a hundred times faster than a CPU. That is not a speed-up; it is a different universe of what is reachable. Every large model since is downstream of this one fact: the hardware for matrix math got cheap and absurdly fast.
Parallelism is the whole trick
Training a network is embarrassingly parallel: every example in a batch can be processed at once, and every neuron in a layer can fire at once. The GPU exists to do many independent things simultaneously — which is exactly what a matrix multiply is.