AI Foundationspredict · compress · act
act XIV

AI for Science

Two chapters of results, one honest chapter about what is left. The pattern that made them work is narrow — and the part of science nobody has automated is the part that decides what to measure.

65

Doing Science

frontieras of 2026-09
before this →Simulating the World

Proteins, forecasts, crystals. It is tempting to generalise from those to “AI does science now”. The rule that actually survived contact with all three is narrower than that, and stating it precisely is more useful than another round of enthusiasm.

THE CLAIM“our model discovered a new drug / material / theorem”1 · Is the output an object?a structure, a candidate, a number —something with a definite value2 · Can an instrument check it?an experiment, a measurement,a competition with hidden answers3 · Was it in the training data?if yes, this is a lookupwearing the costume of a discoveryWHAT SURVIVES THE FILTERSpeed on a known target. Ranking, screening, approximating.AI is excellent at answering questions whose answers are already defined by someone else —and it is not yet good at the step before that, which is deciding which question is worth asking.the filter is also the whole of machine learning in one line: no check, no science.
the filter. Any claim that a model did science can be tested with three questions — and the third one is the one that usually fails.

The failures that are specific to science

It invents the literature. A language model asked to support a claim will produce references that look exactly right and do not exist — authors, journal, year, volume, all plausible. Every citation that arrives from a model has to be checked. This is not a temporary defect of small models; it is what a system optimised for fluent continuation does when it does not know.

It has already seen the answer. Contamination — the benchmark problem from the evaluation chapter — is far more dangerous here, because a scientific benchmark is drawn from the same published literature as the training data. If the protein whose structure you “predicted” was in the training set, you measured recall. Blind competitions like CASP exist precisely because this failure is so easy to commit accidentally.

The reporting gap. Almost every headline of the form “AI discovered X” means: the model proposed X, and a human performed the experiment that found out whether X was true. The discovery is the verification, and the verification is still done by people, slowly, at a bench or in a lab. GNoME’s 380,000 candidates are candidates until someone grows them.

It is not yet an instrument. A result that depends on these weights, this data snapshot and this code version is a claim, not a tool. An instrument is something another lab can pick up and use. That is a much higher bar than a paper.

The part that is not automated

The current ambition is the AI scientist: a system that reads the literature, proposes hypotheses, then designs and runs analyses and writes up what it found. Pieces of that loop are genuinely automatable — searching, extracting, proposing, drafting, even critiquing a proposal.

What is missing is the evaluation step, and the reason is structural. If the same model judges its own hypotheses, nothing has been checked. That is the reward-hacking problem from the alignment act wearing a lab coat: a system that scores itself will learn to produce proposals its own scorer likes, which is not the same as producing proposals that are true.

For a machine to do science it would have to be wrong in a detectable way — to propose things whose failure teaches you something, and to accept the answer. That is the scientific method, and it is not a language skill.

Where that leaves the claims

what to believe, in one place

Believe: proteins, weather, materials — targets with instruments. Believe that prediction is now fast, cheap and sometimes better than the slow method. Believe that the literature-review and analysis assistance is real, and already changing how fast a researcher can work.

Discount: autonomy, and anything that reports a discovery without a name attached to the experiment that confirmed it. Discount “the model reasoned its way to a new law”, and discount any number that came from a benchmark whose contents might have been in the training set.

the honest noteThe uncomfortable summary is that AI has been most useful in science exactly where it is least like a scientist — as a very fast, very literal instrument. That is a real contribution, and it is a different one from the story being told about it.
PREDICTcheap, checkable, working
PROPOSEsometimes useful
ASKnot automated

The end of the line

The guide began with a single bit and a definition: an AI system infers the state of a world and, when it acts, chooses actions toward goals under uncertainty. Sixty-four chapters later, everything in that sentence has been earned.

Information gave us how to represent what we do not know. Probability gave us how to reason about it. Gradients gave us how to learn. Attention gave us the machine. Scaling gave it size, and the learning floor explained why size has the price it does. Alignment gave it manners. Tools and harnesses gave it hands for the digital world. Images, video and worlds gave it imagination, embodiment gave it a body, and this act gave it an instrument.

And through all of it runs one thread, which chapter 2 stated and every layer has confirmed: it is all compression. Entropy is code length, cross-entropy loss is bits per token, a generalisation bound is a bill for capacity, a description length is the size of an explanation, and understanding a protein is finding the short description of a fold.

The unfinished part is the same as it was at the start, and it is one word: uncertainty. Every layer still ends in a guess — the model infers, it does not know. What has changed is the price of being wrong in a simulation, and the price of being wrong in a room. Whether a machine can do the thing that has no checkable answer — choose what to measure next — is the open question, and it is the oldest one here: not can it predict, but can it ask.

introduces →contamination
← previousSimulating the World