The Harness
A model on its own is not an agent. It cannot remember, cannot reach the internet, and cannot run a command. Something has to hold the loop, keep the context, hand it tools, and decide when to stop. That something is the harness, and once you have its name you see it everywhere.
The definition worth remembering
Agent = Model + Harness.
The model holds the capability. The harness holds everything else: the loop, the context, the tools, the memory, the sandbox, the permissions, the orchestration, and the traces. The clean way to say it: the harness owns every decision you can make without retraining the model. That is a large and growing share of what decides whether a system works.
This discipline got its name — harness engineering — as coding agents made the point undeniable. Two products can run the same open-weight model and behave completely differently, because one manages context well and the other does not. Changing the model is expensive and slow; changing the harness is cheap and immediate. Most real progress in agents has come from the second.
The workload that named it
Coding agents are the reason anyone says “harness” out loud. A repository is the ideal environment for an agent: the files are text, the tools are exact, and — crucially — the result can be checked. A coding agent reads the issue, finds the relevant files, edits them, and runs the tests. When a test fails it gets a signal rather than an opinion, and the loop can use it. Benchmarks such as SWE-bench measure exactly this, and what they measure is the system: the same model scores differently under different scaffolds, because the harness decides how much of the repository is shown, how many attempts are allowed, and how a patch is judged. That is the chapter’s lesson in its sharpest form — an agent benchmark evaluates a model inside a harness, so a bare model number from one is half of the thing that was measured.
The parts, and why each exists
Every component here was already a chapter in the agent act. The harness is the insight that they are one thing with one name, and that its design — not the model’s benchmark score — is what you are usually arguing about.
Why the same model scores differently
A benchmark like SWE-bench does not test a model in the abstract. It tests a model inside a harness — a scaffold that decides how the repository is presented, how many attempts are allowed, and what counts as success. Report two numbers from two harnesses and the comparison means nothing. This is why leaderboards disagree, and why “which model is best” is often really “which harness is best”.
The bridge out of text
Almost everything in this guide so far has been text: the model reads text and writes text, and the harness drives that exchange. But text is not the only thing a model can produce or understand. The rest of this act is what lies past it — images first, then models that handle every modality at once, and eventually systems that act in the world. The harness is the bridge: the same idea — a model plus the machinery around it — applies to every one of them.