Memory in a Loop
Some data has no fixed size: a sentence, a song, a stock chart. You cannot treat it as a flat grid. A sequence model reads one element at a time and carries forward a summary of everything it has seen.
The loop
A recurrent neural network (RNN) keeps a hidden state h. At each step it reads an
input, mixes it with the previous state, and produces a new state. The same weights are
reused at every position, so it can handle any length. Training uses
backpropagation through time: unroll the loop and treat each step as a layer.
Why it was hard
Unrolling a thousand steps produces a thousand-layer network, and blame has to travel all the way back. The gradient vanishes or explodes. The LSTM (1997) solved much of this with gates — little learned switches that decide what to pass on. Its first form had input and output gates; a forget gate was added in 2000; together the three became the standard cell. It worked, and for two decades it was the best tool for language.
But there is a deeper problem. Recurrence is inherently sequential: step 5 cannot start until step 4 is done. You cannot parallelise it across a GPU the way you can a convolution. As sequences grew longer and data grew larger, that became fatal.
The sequential bottleneck
An RNN is a long chain of dependencies. A CNN is a wide grid that can be computed all at once. When hardware got parallel and sequences got long, the architecture that looked most natural became the slowest.