Physical AI
The reinforcement chapter described an agent choosing actions toward a goal. It was written abstractly, with bandits and grids, because that is how the idea is easiest to see. Everything since — tokens, attention, scaling, tools, images, video, worlds — has been building the parts you would need to do it for real. This chapter is what happens when the loop closes and the action has to move a physical object.
One model from seeing to moving The architecture that made this work is the vision-language-action model, or VLA: take a model that already understands pictures and instructions, then teach it to emit motor commands as well. Google DeepMind’s RT-2 (2023) is the paper everyone points
at: it showed that a vision-language model trained on internet data could be used directly for robotic control, carrying over general knowledge it was never taught for the task. The same survey literature calls it the point where perception, reasoning and control stopped being three systems and became one.
That is a big deal for a specific reason. Before VLAs, a robot arm needed a policy trained for it. After them, a model that has read about the world can be aimed at a task in words — “put the apple in the bowl” — and generalise to an apple nobody trained it on.
Actions are not tokens
Here is the practical snag. Text is discrete: a vocabulary of 50,000 symbols, chosen one at a time. A joint angle is a real number that changes continuously, and a policy that emits it 50 times a second cannot pick from a list.
Two answers emerged, and they will look familiar. One is action tokenisation: chop continuous movement into discrete chunks so an ordinary language-model head can produce them. The other is to go back to the flow matching and diffusion machinery from the image chapters, which was built precisely for generating continuous things. Physical Intelligence’s π0 does the second — a vision-language model for understanding, plus a separate action expert that produces continuous actions by flow matching. The technique invented to paint pictures turned out to be how you move a robot arm.
The data problem
A language model learns from the whole internet. There is no internet of robot actions. Every demonstration has to be collected by a human teleoperating a machine, slowly, one task at a time — which makes robot data the scarcest resource in the field, and explains why chapter 61 exists at all. If you cannot collect enough real experience, you generate it: a world model is a factory for practice.
The results have started to look like generalisation rather than memorisation. π0.5 (Physical Intelligence, April 2025) can tidy kitchens and bedrooms in houses that were not in its training data, running multi-stage behaviours lasting ten to fifteen minutes. Gemini Robotics 1.5 (DeepMind, September 2025) pairs a VLA with an embodied-reasoning model and — the interesting part — transfers motion across embodiments, so a skill learned on one body helps another. The line has moved on since; the direction has not.
Bodies
Figure’s Helix (February 2025) was the first VLA to output high-rate continuous control of an entire humanoid upper body — wrists, torso, head and individual fingers together, with two robots able to coordinate. Helix 02 (January 2026) extended that to the whole body: walking, manipulating and balancing as one learned system, running a logistics task autonomously for 200 hours. By Helix 2.5 (September 2026) the claim was zero-shot generalisation to thirty homes it had never seen.
Why humanoids, when a wheeled arm is simpler and cheaper? Because the world is already built for that shape — stairs, door handles, shelves, tools. It is not a technical argument, it is a decision to avoid rebuilding civilisation to suit the robot.
Making the jump from simulation
Training in a simulator and deploying in a body is called sim-to-real. It is where the loop leaks. A simulated robot never gets tired, never has a sticky joint, never drops a cup because the table was slightly lower than the model assumed. Physics engines are approximations, so a policy can exploit a simulation’s shortcuts and then fail on hardware — the classic version being a virtual gripper that works fine because the simulator never let anything slip. Solving this is partly better physics (NVIDIA’s Omniverse and Cosmos, the open-source Genesis engine) and partly deliberate randomisation: make the simulation messy on purpose so the real world is one more variation.
This is also why chapter 61’s planner and this chapter’s VLA are converging. If a world model can generate the situation, and a VLA can act in it, then the model of the world and the policy that uses it become two halves of one system — which is what the term “world action model” is starting to mean.
Where this leaves us
Stand back at the end of the line. The guide opened with a definition — an AI system infers the state of a world and, when it acts, chooses actions toward goals under uncertainty — and then spent fifty chapters earning it, one piece at a time. The bits gave us how to represent what we do not know. Probability gave us how to reason about it. Gradients gave us how to learn. Attention gave us the machine. Scaling gave it size. Alignment gave it manners. Tools and harnesses gave it hands for the digital world. Images, video and worlds gave it imagination. And this chapter gave it a body.
What has not been solved is the part the framework named from the start: uncertainty. Every layer still ends in a guess — the model infers, it does not know. A system that acts inside a model can be wrong quickly and cheaply. A system that acts inside a body is wrong in a room, with things in it. That gap, between a prediction and a consequence, is the oldest problem in this history, and the newest one in these machines. It is still open.