Watching the Loop
With a chatbot you can read the output. With an agent, the output may be fine while the process was a disaster — or vice versa. Observability is seeing the entire run: every prompt, tool call, observation, decision, and cost.
What a trace contains
A trace is a tree of spans: the parent run, each model call, each tool invocation, each retry. Good traces record inputs, outputs, latency, token counts, cost, and which model version was used. When something goes wrong three steps back, the trace is the only way to find it. It is also the raw material for improvement: replay the failure, change the prompt, re-run.
The eval loop
An evaluation harness turns real traces into a regression suite. Collect the runs that failed, add them as golden cases, and run them on every change. This is how agent development stops being vibes and becomes engineering. Drift is the ongoing threat: models get updated under you, APIs change, user behaviour shifts. Without a harness you will not notice until a customer does.
Guardrails
Guardrails are checks around the model: input filters, output validators, PII scrubbers, tool-call policies, budget limits. They are not a substitute for alignment. They are the seatbelt — cheap, boring, and they save you when something else fails.
Close the loop
Traces tell you what happened. Evals tell you whether it was good. Together they form the feedback loop that turns a prototype into a system you can trust with real work.