AI Foundationspredict · compress · act
act IX

The Agent Infrastructure

Agents fail quietly. Observability is how you find out what actually happened, step by step.

40

Watching the Loop

before this →A Safe Place to Act

With a chatbot you can read the output. With an agent, the output may be fine while the process was a disaster — or vice versa. Observability is seeing the entire run: every prompt, tool call, observation, decision, and cost.

user turnplantool: searchmodeltool: writeanswertime →
a trace is the agent's run, unfolded — where the real debugging happens

What a trace contains

A trace is a tree of spans: the parent run, each model call, each tool invocation, each retry. Good traces record inputs, outputs, latency, token counts, cost, and which model version was used. When something goes wrong three steps back, the trace is the only way to find it. It is also the raw material for improvement: replay the failure, change the prompt, re-run.

The eval loop

An evaluation harness turns real traces into a regression suite. Collect the runs that failed, add them as golden cases, and run them on every change. This is how agent development stops being vibes and becomes engineering. Drift is the ongoing threat: models get updated under you, APIs change, user behaviour shifts. Without a harness you will not notice until a customer does.

Guardrails

Guardrails are checks around the model: input filters, output validators, PII scrubbers, tool-call policies, budget limits. They are not a substitute for alignment. They are the seatbelt — cheap, boring, and they save you when something else fails.

Close the loop

observability → evaluation → improvement

Traces tell you what happened. Evals tell you whether it was good. Together they form the feedback loop that turns a prototype into a system you can trust with real work.

the disciplineIf you cannot replay a run, you cannot fix it. Instrument first, then optimise.
TRACErecord every step, input, and cost
EVALUATEscore outcomes and trajectories
HARDENguardrails, regressions, retries
end of the stackThat is the full agent infrastructure: a model, tools, memory, protocols, orchestration, serving, isolation, and feedback. One act remains.
introduces →tracingobservabilityevaluation harnessdriftguardrail
← previousA Safe Place to Actnext →Thinking Before Answering