AI Foundationspredict · compress · act
act XI

The Open-Weight World

A small lab in Hangzhou made frontier-adjacent models for a fraction of the expected cost — and put the reasoning weights in public under a permissive licence.

45

DeepSeek

frontieras of 2026-09
before this →What 'Open' Means

DeepSeek is the clearest case study in this act: a lab that treated efficiency as its whole strategy, published the details, and then released the most consequential weights of 2025. To understand why it mattered, you need two ideas from earlier in the guide, used unusually well.

TOKENStoken 1token 2token 3token 4ROUTERpick top-kexpert 1expert 2 ●expert 3expert 4 ●shared expert (always used)only a few experts run per token→ capacity of 671B, cost of 37B
a mixture of experts: each token goes to a few small experts, not one giant network. Big capacity, small active cost.

Two efficiency ideas

Mixture of experts (MoE). A normal network uses all its parameters for every token. An MoE splits the work into many small experts and a router sends each token to only a few of them. Total capacity is huge; the compute per token stays small. DeepSeek’s version adds always-on shared experts and uses very fine-grained experts, so the split is finer than earlier designs.

Multi-head latent attention (MLA). Attention normally stores a growing KV cache as the context gets longer, and that store is what makes serving expensive. MLA compresses the cache into a small latent vector before storing it, so long contexts cost far less memory. It was introduced in DeepSeek-V2 and became a standard trick.

What V3 and R1 actually added

2024-12 · 2025-01

V3 is a 671-billion-parameter MoE that activates 37 billion per token. It trains in FP8 — an 8-bit number format — which cuts memory and speeds the maths, and it balances expert load without an extra loss term (auxiliary-loss-free load balancing). It also predicts several future tokens at once (multi-token prediction). R1 is the reasoning model: trained with GRPO, a reinforcement method that scores a group of answers against each other instead of learning a separate value function.

the surprising resultR1-Zero was trained by reinforcement learning with no supervised examples at all — and long chains of reasoning appeared on their own. It was hard to read, so the final R1 added a small amount of clean starting data.
V2MLA + DeepSeekMoE make it cheap to run
V3FP8 + better routing make it cheap to train
R1GRPO makes it reason — released under MIT
V4hybrid attention + 1M context (2026)

The next turn: V4

By April 2026 the same lab had moved again. DeepSeek-V4 arrived as a preview in two sizes — V4-Pro at 1.6 trillion parameters (49B active per token) and V4-Flash at 284B (13B active) — both holding a one-million-token context. The architectural news is that V4 retires the very trick described above: the MLA compression that made V2 cheap is replaced by a hybrid attention scheme that mixes Compressed Sparse Attention with Heavily Compressed Attention, and ordinary residual connections give way to manifold-constrained hyper-connections. The pattern is the point. Each generation buys the same thing with a new mechanism — more context, at less cost per token — and the mechanism that was the frontier last year becomes the thing being replaced.

Why it shook things

Two things landed at once. First, the cost: DeepSeek published the compute behind V3, and the figure was far below what a frontier model was assumed to need. When R1’s app topped the app stores in January 2025, the market read it as evidence that the huge capex forecasts might be wrong, and chip stocks fell sharply.

Second, the licence. R1’s weights went out under MIT — one of the most permissive terms that exists — together with distilled versions trained to imitate R1’s reasoning at small sizes (knowledge distillation). That let anyone run a reasoning model locally. The next chapter maps the rest of the field that formed around this moment.

introduces →mixture of expertsmulti-head latent attentionauxiliary-loss-free load balancingmulti-token predictionGRPOFP8knowledge distillation
← previousWhat 'Open' Meansnext →The Chinese Labs