AI Foundationspredict · compress · act

concepts → Model

Model

RT-2

frontieras of 2026-092 connections

the paper that aimed a vision-language model at a robot arm and let the web knowledge carry over

Where it sits

The prism has six jobs across and eight layers down. Its primary cell is no mode × L1. Hatched cells cannot exist — a GPU does not learn, an institution does not infer.

Evidence: benchmark.

What must come first — and what it unlocks

Left to right is reading order, derived from the prerequisite_of edges. Nothing here is hand-ordered: the diagram is the graph.

nothing comes first — this is a starting pointRT-2nothing depends on it yet — a leaf in the reading order

Where to read it

The chapter that introduces it, and any chapter that uses it again.

62Physical AIact XIII · Beyond Text

Where it comes from

Every connection

All 2 edges touching this node, grouped by relation family — the sections above are highlights from this list. Colours match the relation families inthe atlas.

Structure · 1
is an instance ofVision-language-actionArchitecture
Flow · 1
usesLanguage modelArchitecture

This page is a projection of one node in src/data/concepts.ts. It has no prose file of its own — 521 declared edges produce all 340 of these pages. Edit an edge and both endpoints change.