Doing Science
Proteins, forecasts, crystals. It is tempting to generalise from those to “AI does science now”. The rule that actually survived contact with all three is narrower than that, and stating it precisely is more useful than another round of enthusiasm.
The failures that are specific to science
It invents the literature. A language model asked to support a claim will produce references that look exactly right and do not exist — authors, journal, year, volume, all plausible. Every citation that arrives from a model has to be checked. This is not a temporary defect of small models; it is what a system optimised for fluent continuation does when it does not know.
It has already seen the answer. Contamination — the benchmark problem from the evaluation chapter — is far more dangerous here, because a scientific benchmark is drawn from the same published literature as the training data. If the protein whose structure you “predicted” was in the training set, you measured recall. Blind competitions like CASP exist precisely because this failure is so easy to commit accidentally.
The reporting gap. Almost every headline of the form “AI discovered X” means: the model proposed X, and a human performed the experiment that found out whether X was true. The discovery is the verification, and the verification is still done by people, slowly, at a bench or in a lab. GNoME’s 380,000 candidates are candidates until someone grows them.
It is not yet an instrument. A result that depends on these weights, this data snapshot and this code version is a claim, not a tool. An instrument is something another lab can pick up and use. That is a much higher bar than a paper.
The part that is not automated
The current ambition is the AI scientist: a system that reads the literature, proposes hypotheses, then designs and runs analyses and writes up what it found. Pieces of that loop are genuinely automatable — searching, extracting, proposing, drafting, even critiquing a proposal.
What is missing is the evaluation step, and the reason is structural. If the same model judges its own hypotheses, nothing has been checked. That is the reward-hacking problem from the alignment act wearing a lab coat: a system that scores itself will learn to produce proposals its own scorer likes, which is not the same as producing proposals that are true.
For a machine to do science it would have to be wrong in a detectable way — to propose things whose failure teaches you something, and to accept the answer. That is the scientific method, and it is not a language skill.
Where that leaves the claims
Believe: proteins, weather, materials — targets with instruments. Believe that prediction is now fast, cheap and sometimes better than the slow method. Believe that the literature-review and analysis assistance is real, and already changing how fast a researcher can work.
Discount: autonomy, and anything that reports a discovery without a name attached to the experiment that confirmed it. Discount “the model reasoned its way to a new law”, and discount any number that came from a benchmark whose contents might have been in the training set.
The end of the line
The guide began with a single bit and a definition: an AI system infers the state of a world and, when it acts, chooses actions toward goals under uncertainty. Sixty-four chapters later, everything in that sentence has been earned.
Information gave us how to represent what we do not know. Probability gave us how to reason about it. Gradients gave us how to learn. Attention gave us the machine. Scaling gave it size, and the learning floor explained why size has the price it does. Alignment gave it manners. Tools and harnesses gave it hands for the digital world. Images, video and worlds gave it imagination, embodiment gave it a body, and this act gave it an instrument.
And through all of it runs one thread, which chapter 2 stated and every layer has confirmed: it is all compression. Entropy is code length, cross-entropy loss is bits per token, a generalisation bound is a bill for capacity, a description length is the size of an explanation, and understanding a protein is finding the short description of a fold.
The unfinished part is the same as it was at the start, and it is one word: uncertainty. Every layer still ends in a guess — the model infers, it does not know. What has changed is the price of being wrong in a simulation, and the price of being wrong in a room. Whether a machine can do the thing that has no checkable answer — choose what to measure next — is the open question, and it is the oldest one here: not can it predict, but can it ask.