Thinking Before Answering
For years, the recipe was: train a huge model, then read off its first answer. The frontier shifted. Now a model can think before it answers — generating a long internal chain of reasoning, checking itself, and only then committing.
Chain-of-thought
Chain-of-thought (CoT) is startlingly simple: ask the model to reason step by step before giving the answer, and accuracy rises on maths, logic, and multi-step problems. The tokens of reasoning give the model working memory — it can hold intermediate results that would otherwise have to fit in a single forward pass.
Reasoning models
A reasoning model is trained specifically to use a long internal thought process: reinforcement learning rewards correct final answers and lets the model discover productive reasoning patterns. At inference, it may generate thousands of hidden tokens, backtrack, try another approach, and verify. Test-time compute or inference scaling is the knob: how much thinking to buy. Some systems even sample many answers and take a majority vote, or search over reasoning paths.
The new scaling axis
This reframes the economics. A smaller model that thinks for longer can beat a bigger model that answers immediately — and you only pay for the hard questions. Compute becomes a runtime dial instead of a fixed architecture.