Rolling Downhill
A model has millions of numbers — parameters — that it can tune. A loss function scores how wrong it is. The question is: which way should each number move to reduce the loss? The answer is the gradient.
Gradient descent in one sentence
Compute the slope of the loss with respect to every parameter, then nudge each parameter a little in the opposite direction.
The learning rate controls the nudge. Too small and training crawls; too large and you leap over the valley and diverge. Finding a good rate — and a schedule that shrinks it over time — is one of the most consequential practical choices in all of deep learning.
Backpropagation: the gradient, cheaply
For a deep network, computing every slope by hand would be hopeless. Backpropagation does it in one sweep: the chain rule, applied from the output layer back to the input, reusing each intermediate result. It is the reason deep learning is possible at all. Forward pass: make a prediction. Backward pass: assign blame to every parameter.
The engine of the whole field
Every model in the rest of this guide — CNNs, RNNs, transformers, diffusion — trains with this same loop. Architectures differ; the optimizer barely does.