The Open Question
Look back at the shape of the journey. We started with a bit of uncertainty, and we ended with a system that perceives, remembers, reasons, calls tools, and acts under constraint. Every step reused the step before it. Nothing appeared from nowhere.
Three things we now know
Compression is the engine. The through-line from Shannon to a language model is unbroken. A model that predicts well has found structure; finding structure is compressing; compressing is understanding.
Scale changes behaviour, not just scores. Straight lines on log-log plots turned capability into a forecast and made compute the strategic resource. The bitter lesson kept winning, but inductive bias still decides how efficiently you learn.
Infrastructure decides what ships. The last act of this guide was almost entirely engineering: memory, protocols, sandboxes, serving, tracing. Capability without that layer stays a demo.
Three things we do not know
Is scale enough? The scaling debate is unresolved. One camp expects continued smooth gains and eventual generality; the other expects diminishing returns and missing pieces. Both have been right before.
Will alignment keep up? Oversight scales with human attention, capability scales with compute. Closing that gap — through interpretability, evaluations, and control systems — is the central technical and political problem of the field.
What are we building? AGI and superintelligence are words with no agreed definitions and enormous rhetorical weight. It is worth remembering that the same phrases were used with confidence in 1965 and 1985. Caution runs in both directions: against dismissing the technology, and against believing its own marketing.
Where to go from here
You now hold the whole line: a bit, an entropy, a gradient, a perceptron, an attention head, a tool call, a sandbox, a trace. Each is simple. Their composition is not. That gap — between simple parts and complex behaviour — is the whole subject.