AI Foundationspredict · compress · act
act IX

The Agent Infrastructure

Complex work needs more than one agent step. Orchestration is how you structure many steps — and how you make them survive failure.

37

Many Hands, One Job

before this →A Port for Tools

One model call rarely finishes a real task. You need a structure: who does what, in what order, and what happens when a step fails. That structure is orchestration.

CHAINROUTERSUPERVISORleaddelegate & verify
three common shapes: chain, router, and supervisor with workers

Three patterns

A workflow is a fixed path: step A then step B. Deterministic, easy to reason about, and often the right answer — most “agents” should be workflows. A router sends each request to a specialist. A supervisor (or multi-agent) pattern has a lead model that delegates to workers and checks their results. More autonomy buys flexibility and costs predictability; reach for it only when the path genuinely branches.

Durable execution

Agents are long-running, so they fail mid-task: a network error, a rate limit, a process restart. Durable execution treats an agent run as a resumable state machine. After every step, the state is saved — checkpointing — so a crash resumes from the last good step instead of starting over. This is ordinary distributed-systems engineering, and it is what separates a demo from something you can leave running.

Autonomy is a cost

choose the simplest structure that works

Every extra agent step adds latency, cost, and failure modes. A two-step workflow with a good prompt beats a five-agent swarm with a vague one. Frameworks should make the simple case easy, not the elaborate case tempting.

the principleStructure follows the task. Deterministic where you can be, agentic only where you must be.
PLANdecompose the job
EXECUTErun steps, checkpoint each one
RECOVERresume, retry, or escalate
introduces →orchestrationworkflowmulti-agentdurable executioncheckpointing
← previousA Port for Toolsnext →How Tokens Get Served