AI Foundationspredict · compress · act

concepts → Architecture

Architecture

Transformer

decade7 connections

attention and feed-forward blocks stacked — the architecture that swallowed the field

Where it sits

The prism has six jobs across and eight layers down. Its primary cell isrepresent × L1. Hatched cells cannot exist — a GPU does not learn, an institution does not infer.

What must come first — and what it unlocks

Left to right is reading order, derived from the prerequisite_of edges. Nothing here is hand-ordered: the diagram is the graph.

AttentionTransformerMixture of expertsPretraining

Before it: Attention

Unlocks: Mixture of experts · Pretraining

What kind of thing it is — and what it is made of

It is made of: Attention.

Where to read it

The chapter that introduces it, and any chapter that uses it again.

24Attention Is All You Needact 6 · The Attention Revolution

Where it comes from

paperAttention Is All You NeedAshish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, Illia Polosukhin · 2017

Every connection

All 7 edges touching this node, grouped by relation family — the sections above are highlights from this list. Colours match the relation families inthe atlas.

Structure · 1
has as a partAttentionMechanism
Order · 3
requiresAttentionMechanism
must come beforeMixture of expertsMechanism
must come beforePretrainingAlgorithm
Flow · 1
is used byDiffusion transformerArchitecture
Lineage · 1
improves onRecurrent neural networkArchitecture
Contrast · 1
trades off withLSTMArchitecture

This page is a projection of one node in src/data/concepts.ts. It has no prose file of its own — 230 declared edges produce all 236 of these pages. Edit an edge and both endpoints change.