AI Foundationspredict · compress · act
act XII

Beyond Text

An agent is a model plus everything that is not the model. That second part now has a name, and it decides most of what actually works.

52

The Harness

decadeas of 2026-09
before this →Watching the Loop

A model on its own is not an agent. It cannot remember, cannot reach the internet, and cannot run a command. Something has to hold the loop, keep the context, hand it tools, and decide when to stop. That something is the harness, and once you have its name you see it everywhere.

THE MODELweights · fixedloop policywhen to act, when to stopcontext windowwhat to keep, what to droptoolsschemas and resultsmemorynotes, files, retrievalsandboxwhere actions happenpermissionswhat it may never doorchestrationother agents, other stepstelemetrytraces, cost, failures
the harness is everything around the weights. The model supplies the intelligence; the harness decides what it sees, what it can do, and how long it keeps trying.

The definition worth remembering

Agent = Model + Harness.

The model holds the capability. The harness holds everything else: the loop, the context, the tools, the memory, the sandbox, the permissions, the orchestration, and the traces. The clean way to say it: the harness owns every decision you can make without retraining the model. That is a large and growing share of what decides whether a system works.

This discipline got its name — harness engineering — as coding agents made the point undeniable. Two products can run the same open-weight model and behave completely differently, because one manages context well and the other does not. Changing the model is expensive and slow; changing the harness is cheap and immediate. Most real progress in agents has come from the second.

The parts, and why each exists

loop
Loop policy — the ReAct cycle from the agent chapter: think, act, observe, repeat — and the rule for when to stop.
ctx
Context management — the harness chooses what enters the window and what is evicted. This is the single biggest lever on quality and cost.
tools
Tool schemas — each tool's name, arguments and result format. Vague schemas cause most agent mistakes.
mem
Memory — what survives one turn: scratch files, notes, and retrieval from outside the weights.
sand
Sandbox — the place actions actually happen, isolated so a wrong step is survivable.
perm
Permissions — the rules the agent may never cross — the safety boundary, not a prompt.
orch
Orchestration — multiple agents, retries, and durable steps that survive a crash.
obs
Telemetry — traces, token counts, and failures — without which none of the above can be improved.

Every component here was already a chapter in the agent act. The harness is the insight that they are one thing with one name, and that its design — not the model’s benchmark score — is what you are usually arguing about.

Why the same model scores differently

the evaluation scaffold

A benchmark like SWE-bench does not test a model in the abstract. It tests a model inside a harness — a scaffold that decides how the repository is presented, how many attempts are allowed, and what counts as success. Report two numbers from two harnesses and the comparison means nothing. This is why leaderboards disagree, and why “which model is best” is often really “which harness is best”.

the practical ruleWhen an agent fails, look at the harness before the model. Most failures are a missing tool, a vague schema, a context that dropped the wrong thing, or a loop with no stopping rule.
MODELfixed capability
HARNESScheap to change, decides the result
AGENTwhat actually ships

The bridge out of text

Almost everything in this guide so far has been text: the model reads text and writes text, and the harness drives that exchange. But text is not the only thing a model can produce or understand. The rest of this act is what lies past it — images first, then models that handle every modality at once, and eventually systems that act in the world. The harness is the bridge: the same idea — a model plus the machinery around it — applies to every one of them.

introduces →agent harnessharness engineeringcontext managementloop policyscaffolding
← previousThe Open Frontiernext →Image Models