Attention Is All You Need
The bottleneck was sequence. RNNs had to read one token at a time. The transformer (2017) asks a different question: what if every token could look directly at every other token, all at once?
Queries, keys, values
Each token produces three vectors. A query is what it is looking for. A key is what it offers. A value is what it will hand over. Attention scores every query against every key, softmaxes those scores into weights, and returns a weighted sum of the values. In short:
Every token takes a weighted average of the information it finds relevant.
This is self-attention, and its magic is that the path between any two tokens is exactly one step. No vanishing gradient across distance. No sequential loop. The whole sequence is processed in parallel as a matrix multiply — perfect for GPUs.
Because attention itself has no notion of order, we inject positional encodings, so the model knows that “dog bites man” differs from “man bites dog.” Modern variants (RoPE, relative positions) handle this more gracefully, but the principle is the same.
The block
A transformer is a stack of identical blocks. Each block: self-attention to mix information across tokens, then a small feed-forward network to process each token independently, each wrapped in a residual connection and normalisation. That is the whole architecture. Scale it up and it reads the internet.