concepts → Mechanism
Tokenizer
the rule that cuts text into tokens — it shapes everything downstream
Where it sits
The prism has six jobs across and eight layers down. Its primary cell isrepresent × L1, with a noted second cell. Hatched cells cannot exist — a GPU does not learn, an institution does not infer.
represent×L2 — tokenization is a data-layer choice, made before training
What must come first — and what it unlocks
Left to right is reading order, derived from the prerequisite_of edges. Nothing here is hand-ordered: the diagram is the graph.
Unlocks: Pretraining
What kind of thing it is — and what it is made of
Examples of it: Byte-pair encoding.
It is made of: Token, Vocabulary, Speech token.
Where to read it
The chapter that introduces it, and any chapter that uses it again.
Where it comes from
Every connection
All 5 edges touching this node, grouped by relation family — the sections above are highlights from this list. Colours match the relation families inthe atlas.
This page is a projection of one node in src/data/concepts.ts. It has no prose file of its own — 230 declared edges produce all 236 of these pages. Edit an edge and both endpoints change.