Learning from Blame
A single neuron can only draw a straight line. Stack neurons in layers and the network can bend, curve and carve out any shape. But for years, no one had a practical way to train the hidden layers — the neurons in the middle, whose correct outputs nobody knows.
The algorithm
Backpropagation (popularised in 1986 by Rumelhart, Hinton, and Williams) solves the credit-assignment problem. Make a prediction, measure the error, then send the blame backward through the network using the chain rule. Each weight learns how much it contributed to the mistake, and moves to reduce it.
A multilayer perceptron — an input layer, one or more hidden layers, an output layer — is the result. With enough hidden units it can approximate essentially any function. The universal approximation theorem made it official.
What went wrong anyway
Blame weakens as it travels backward. In a deep network, the signal shrinks at every layer until early layers learn almost nothing — the vanishing gradient problem. This is why “deep” learning had to wait: it took better activations (ReLU), better initialisation, and later skip connections to let the gradient flow.