AI Foundationspredict · compress · act

concepts → Mechanism

Mechanism

Mixture of experts

also called MoE

decade4 connections

route each token to a few small experts — huge capacity, small active cost

Where it sits

The prism has six jobs across and eight layers down. Its primary cell isrepresent × L1, with a noted second cell. Hatched cells cannot exist — a GPU does not learn, an institution does not infer.

Overlays: Efficiency, cost & energy. learn×L4 — balanced routing is a training-time problem (the auxiliary loss)

What must come first — and what it unlocks

Left to right is reading order, derived from the prerequisite_of edges. Nothing here is hand-ordered: the diagram is the graph.

TransformerMixture of expertsDeepSeek-V3

Before it: Transformer

Unlocks: DeepSeek-V3

What kind of thing it is — and what it is made of

It is made of: Auxiliary-loss-free load balancing.

Where to read it

The chapter that introduces it, and any chapter that uses it again.

45DeepSeekact 10 · The Open-Weight World

Where it comes from

paperDeepSeek-V3 Technical ReportDeepSeek-AI · 2024

Every connection

All 4 edges touching this node, grouped by relation family — the sections above are highlights from this list. Colours match the relation families inthe atlas.

Structure · 1
Order · 2
requiresTransformerArchitecture
must come beforeDeepSeek-V3Model
Flow · 1
is used byDeepSeek-V3Model

This page is a projection of one node in src/data/concepts.ts. It has no prose file of its own — 230 declared edges produce all 236 of these pages. Edit an edge and both endpoints change.