AI Foundationspredict · compress · act

concepts → Architecture

Architecture

Vision-language-action

frontieras of 2026-099 connections

one model that turns what it sees and what you asked for into motor commands

Where it sits

The prism has six jobs across and eight layers down. Its primary cell is act × L1. Hatched cells cannot exist — a GPU does not learn, an institution does not infer.

Runs at: datacenter. Evidence: benchmark.

What must come first — and what it unlocks

Left to right is reading order, derived from the prerequisite_of edges. Nothing here is hand-ordered: the diagram is the graph.

nothing comes first — this is a starting pointVision-language-actionnothing depends on it yet — a leaf in the reading order

What kind of thing it is — and what it is made of

Examples of it: RT-2, π0, Helix, Gemini Robotics, OpenVLA.

Where to read it

The chapter that introduces it, and any chapter that uses it again.

62Physical AIact XIII · Beyond Text

Where it comes from

Every connection

All 9 edges touching this node, grouped by relation family — the sections above are highlights from this list. Colours match the relation families inthe atlas.

Structure · 6
is an instance ofPlannerMechanism
has as an instanceRT-2Model
has as an instanceπ0Model
has as an instanceHelixSystem
has as an instanceGemini RoboticsSystem
has as an instanceOpenVLAModel
Flow · 3
usesAction tokenisationMechanism
usesFlow matchingMechanism
is used byHumanoidHardware

This page is a projection of one node in src/data/concepts.ts. It has no prose file of its own — 521 declared edges produce all 340 of these pages. Edit an edge and both endpoints change.