The Compute Squeeze
Every model in this act is a pile of arithmetic that has to run somewhere. For a decade that “somewhere” has meant Nvidia, and starting in 2022 the United States restricted which Nvidia chips could be sold to China. The result is the clearest natural experiment in the field: what happens when a capable group is denied the fastest hardware.
The chokepoint
Export controls are rules that limit the sale of certain technologies across borders. From October 2022 the United States restricted exports of the most advanced AI chips to China, and tightened the rules in 2023 and after. When a specific chip is banned, the seller often makes a slower export version to stay legal — which is exactly what happened, and which the next round of rules then banned too. The rules kept chasing the workarounds.
The practical consequence: Chinese labs trained on chips that were legal to buy, with restricted interconnect — the links between chips — which caps how large a single training run can be. You cannot copy your way out of that. You can only be cleverer about the maths.
The rules can loosen too
Controls are policy, not physics, and policy moves in both directions. After tightening through 2022–2024, the United States partially reversed course: a rule effective 15 January 2026 moved licences for Nvidia’s H200, AMD’s MI325X and comparable chips from a presumption of denial to case-by-case review, subject to security and availability conditions. The wall that shaped this chapter’s models was therefore a moving constraint rather than a settled fact. The models it produced keep their efficiency advantages either way — those turned out to be good engineering, not merely workarounds for a shortage.
The domestic response
The obvious answer is to build your own chips, and China has: Huawei Ascend GPUs and their software stack CANN are the most serious non-American effort. The hard part was never the silicon alone but the software around it — the compilers and libraries that make a chip usable. That gap is measured in years, and closing it is as much a software problem as a hardware one.
Constraint as a design pressure
DeepSeek’s efficiency tricks were not only cleverness for its own sake. MLA shrinks the memory a long context needs. FP8 halves the bytes moved per number. Mixture of experts gives large capacity at small active cost. Each one directly attacks a hardware limit — bandwidth, memory, and compute — that a team with unlimited top-tier chips would have felt less pressure to fix. The squeeze made the field invent things it otherwise might not have.
Everyone’s constraint
The squeeze is geopolitical, but the underlying physics is universal. Most people will never own a cluster, and the model that fits on their machine is the one they can actually use. That is why the next chapters turn from who builds models to how anyone runs one — and what the licence on the file allows them to do with it.