Learning from Blame
A single neuron can only draw a straight line. Stack neurons in layers and the network can bend, curve, and carve out any shape. But for thirty years, nobody could work out how to train the hidden layers — the neurons in the middle, whose correct outputs nobody knows.
The algorithm
Backpropagation (popularised in 1986 by Rumelhart, Hinton, and Williams) solves the credit-assignment problem. Make a prediction, measure the error, then send the blame backward through the network using the chain rule. Each weight learns how much it contributed to the mistake, and moves to reduce it.
A multilayer perceptron — an input layer, one or more hidden layers, an output layer — is the result. With enough hidden units it can approximate essentially any function. The universal approximation theorem made it official.
What went wrong anyway
Blame weakens as it travels backward. In a deep network, the signal shrinks at every layer until early layers learn almost nothing — the vanishing gradient problem. This is why “deep” learning had to wait: it took better activations (ReLU), better initialisation, and later skip connections to let the gradient flow.