AI Foundationspredict · compress · act
act XI

The Open-Weight World

Open weights did not start in China. Llama made them respectable, Mistral proved a small lab could ship them, and a few projects went all the way to fully reproducible.

47

The Other Open Lineages

frontieras of 2026-09
before this →What 'Open' Means

Long before DeepSeek, the open-weights habit was set in the West. Meta’s Llama family made releasing weights a mainstream strategy in 2023, a small French lab called Mistral showed that a handful of people could ship competitive open models, and a research project called OLMo went further than anyone — releasing the data.

2022BLOOM — a 176B model trained in the open, multilingual
2023Llama 2 — open weights go mainstream
2023Mistral 7B, then Mixtral — small lab, strong open MoE
2024OLMo — weights, code and data released
2024Llama 3, Gemma — the giants join in
2025Llama 4 — open mixture-of-experts, natively multimodal
2025Frontier-adjacent open weights become routine

The West’s open line

Llama (Meta) set the template: release weights, allow broad use, keep a size cap. Its licence limits commercial use above a large user threshold and requires a naming credit — openness with a leash. Llama 4 (April 2025) changed the shape of the line rather than its terms: Scout and Maverick are mixture-of-experts models, with far more total parameters than they activate per token, and they are natively multimodal — released under the same community licence, not a permissive one. Mistral (France) went the other way, shipping its flagship models under Apache 2.0, with no cap at all. Gemma (Google) and Phi (Microsoft) followed with capable small models under custom terms. Falcon (TII, United Arab Emirates), Command-R (Cohere), DBRX (Databricks), Granite (IBM), Grok (xAI), and SmolLM (Hugging Face) filled in the rest.

The fully open ones

A separate strand climbs the whole ladder of the last chapter:

What a fully open model actually gives you

OLMo · Pythia · BLOOM

OLMo (Allen Institute for AI) releases the weights, the training code, and the data — every rung, so a stranger can reproduce the run. Pythia (EleutherAI) does the same at small sizes for research. BLOOM (BigScience) trained a 176-billion-parameter multilingual model in public. These are not the strongest models in the world. Their value is different: you can study them — trace a behaviour back to the data that caused it, and run controlled experiments that are impossible on a model you can only call.

why this matters for the whole fieldAlmost all safety and interpretability research depends on being able to see inside a training run. A fully open model is the only kind you can do that with.
WEIGHTS ONLYuse it, measure it, cannot explain it
FULLY OPENrebuild it, ablate it, explain it

Why this line matters

Even the leashed versions changed the floor. Once a capable model is free to download, it becomes an open-weight baseline: the thing every new claim is measured against, and the thing a thousand startups build on instead of paying an API. The competitive effect is the same whether the lab is generous or strategic — the middle of the market gets cheaper every time a strong model is let go.

introduces →LlamaMistralOLMofully open modelopen-weight baseline
← previousThe Chinese Labsnext →The Compute Squeeze