AI Foundationspredict · compress · act
act VII

The Attention Revolution

The 2017 paper that dropped recurrence entirely and let every token look at every other token at once.

24

Attention Is All You Need

before this →Learning from Reward
2014Attention — Bahdanau gives translation a way to look back at the input — the seed of the transformer.
2017The transformer — "Attention Is All You Need" replaces recurrence with attention.

The bottleneck was sequence. RNNs had to read one token at a time. The transformer (2017) asks a different question: what if every token could look directly at every other token, all at once?

querykey 1key 2key 3key 4weightshighsoftmax → weighted sum of values
one query attends to every key; the weighted values become the new representation

Queries, keys, values

Each token produces three vectors. A query is what it is looking for. A key is what it offers. A value is what it will hand over. Attention scores every query against every key, softmaxes those scores into weights, and returns a weighted sum of the values. In short:

Every token takes a weighted average of the information it finds relevant.

This is self-attention, and its magic is that the path between any two tokens is exactly one step. No vanishing gradient across distance. No sequential loop. The whole sequence is processed in parallel as a matrix multiply — perfect for GPUs.

Because attention itself has no notion of order, we inject positional encodings, so the model knows that “dog bites man” differs from “man bites dog.” Modern variants (RoPE, relative positions) handle this more gracefully, but the principle is the same.

The block

what a transformer layer actually contains

A transformer is a stack of identical blocks. Each block: self-attention to mix information across tokens, then a small feed-forward network to process each token independently, each wrapped in a residual connection and normalisation. That is the whole architecture. Scale it up and it reads the internet.

the shiftRecurrence and convolution both encode assumptions about order and locality. Attention encodes almost none — which makes it general, data-hungry, and able to learn structure instead of being told it.
Q · Kᵀscore every pair of tokens
SOFTMAXturn scores into weights
× Vmix the information — done
introduces →attentiontransformerself-attentionquery-key-valuepositional encoding
← previousLearning from Rewardnext →Breaking Language into Pieces