AI Foundationspredict · compress · act
EVERYTHING THAT LED HERE

How we got here

The ideas did not arrive in order, and twice the money ran out. This is the whole arc on one rail — so the winters, the turns, and the sudden acceleration make sense as one story.

1936 → 2026 · 54 moments · 9 eras

settled — will not changeapproach — review in ~5 yearsmoving — review quarterly
era 11936–1950

Before it had a name

The ideas arrive before the field has a word for itself.

era 21956–1973

The first spring

Optimism, symbols, and the first machine that learns.

1958
The perceptronsettled

Rosenblatt builds a machine that learns to tell two classes apart by changing its weights.

1966
ELIZAsettled

Weizenbaum's chatbot reflects your words back — and people confide in it anyway.

era 31980–1987

The first winter, and the expert boom

Funding collapses — then rule-based systems make real money.

1982
The Fifth Generationapproach

Japan bets big public money on parallel logic machines.

era 41988–1997

The second winter, and the statistical turn

The boom busts again, and quiet statistical methods take over.

1995
Support vector machinesapproach

Vapnik's maximum-margin method makes small-data learning rigorous.

1997
Deep Bluesettled

IBM's machine beats world chess champion Garry Kasparov.

era 52006–2009

The quiet run-up

Cheap parallel chips, big data, and one labelled image set.

era 62012–2017

Deep learning arrives

Networks finally work, and the field changes shape.

2014
GANsapproach

Goodfellow pits a generator against a discriminator until the fake looks real.

2015
ResNetapproach

Skip connections let networks get very deep without falling apart.

2016
AlphaGoapproach

A search-plus-network agent beats Go champion Lee Sedol.

era 72018–2021

The transformer era

One architecture, trained bigger, swallows everything.

2018
BERT and GPT-1approach

Pre-train on everything, then fine-tune — the default recipe for years.

2019
GPT-2approach

A large language model writes fluent text; the release is staged.

2020
Scaling lawsapproach

Loss falls as a smooth power law in compute, data and parameters — so you can plan ahead.

2020
GPT-3approach

175 billion parameters, and prompting alone works.

2021
CLIPapproach

Images and text land in one shared space, so you can search one with the other.

era 82022–2024

Products, alignment, prizes

The technology becomes a product, and a public argument.

2022-03
Chinchillaapproach

Most giant models were badly under-trained on data. Fix that, and smaller wins.

2022-11
ChatGPTapproach

Preference tuning turns a text model into a product. The public arrives.

2023
GPT-4moving

Frontier models pass many human exams, and the race goes public.

2023
Open weights get strongmoving

Llama 2 and Mistral make good models runnable and modifiable by anyone.

2024
A Nobel for AIsettled

Hopfield and Hinton share the physics prize; Hassabis and Jumper the chemistry prize.

era 92024–2026

Open weights and reasoning

Strong open models, models that think, and agents that act.

2024-12
DeepSeek-V3moving

A strong open-weight model trained at a fraction of the assumed cost.

2025-01
DeepSeek-R1moving

Open reasoning weights land, with a cost shock and a new rivalry in the open.

2025
The open-weights wavemoving

Qwen, GLM, Kimi and others ship frontier-adjacent models anyone can download.

2026-01
The rules loosenmoving

The US moves advanced-chip licences for China to case-by-case review, easing the squeeze that shaped the efficiency wave.

2026-04
DeepSeek-V4moving

A one-million-token context arrives with hybrid attention, replacing the very MLA trick the V2 line was built on.

Nothing above is a new claim: the settled rows are textbook, the moving rows are dated and will be revised. Colour marks the turning-away —violet for a setback, cyanfor progress. Read alongside the cross-index for the ideas, not the dates.