AI Foundationspredict · compress · act
act III

The Instruments

Learning has one algorithm at its heart: measure the error, then take a small step downhill.

10

Rolling Downhill

before this →Everything Is a Vector

A model has millions of numbers — parameters — that it can tune. A loss function scores how wrong it is. The question is: which way should each number move to reduce the loss? The answer is the gradient.

lossparametersstep = −learning rate × slopethe bottom — where loss is smallest
the gradient points uphill, so we step the other way

Gradient descent in one sentence

Compute the slope of the loss with respect to every parameter, then nudge each parameter a little in the opposite direction.

The learning rate controls the nudge. Too small and training crawls; too large and you leap over the valley and diverge. Finding a good rate — and a schedule that shrinks it over time — is one of the most consequential practical choices in all of deep learning.

Backpropagation: the gradient, cheaply

For a deep network, computing every slope by hand would be hopeless. Backpropagation does it in one sweep: the chain rule, applied from the output layer back to the input, reusing each intermediate result. It is the reason deep learning is possible at all. Forward pass: make a prediction. Backward pass: assign blame to every parameter.

The engine of the whole field

all of deep learning, in one loop

Every model in the rest of this guide — CNNs, RNNs, transformers, diffusion — trains with this same loop. Architectures differ; the optimizer barely does.

the caveatGradient descent finds a low point, not the lowest. The landscape is bumpy and huge. In practice, “good enough” valleys are plentiful and reachable.
FORWARDpredict, and measure the loss
BACKWARDpropagate blame along every weight
UPDATEstep parameters downhill
introduces →loss functiongradient descentbackpropagationlearning rate
← previousEverything Is a Vectornext →Why Memorizing Fails