Memory in a Loop
Some data has no fixed size: a sentence, a song, a stock chart. You cannot treat it as a flat grid. A sequence model reads one element at a time and carries forward a summary of everything it has seen.
The loop
A recurrent neural network (RNN) keeps a hidden state h. At each step it reads an
input, mixes it with the previous state, and produces a new state. The same weights are
reused at every position, so it can handle any length. Training uses
backpropagation through time: unroll the loop and treat each step as a layer.
Why it was hard
Unrolling a thousand steps produces a thousand-layer network, and blame has to travel all the way back. The gradient vanishes or explodes. The LSTM (1997) solved much of this with gates — little learned switches that decide what to remember, what to forget, and what to output. It worked, and for two decades it was the best tool for language.
But there is a deeper problem. Recurrence is inherently sequential: step 5 cannot start until step 4 is done. You cannot parallelise it across a GPU the way you can a convolution. As sequences grew longer and data grew larger, that became fatal.
The sequential bottleneck
An RNN is a long chain of dependencies. A CNN is a wide grid that can be computed all at once. When hardware got parallel and sequences got long, the architecture that looked most natural became the slowest.