How Do We Even Know?
You cannot improve what you cannot measure, and in AI it is dangerously easy to measure the wrong thing. Evaluation is the craft of finding out what a model can really do.
Benchmarks and their lies
A benchmark is a fixed test everyone runs. It is useful and it is inevitably gamed. Once a benchmark becomes a target, models are trained on data that resembles it, and the score stops predicting real ability. Contamination is the acute version: test questions leak into the training corpus, and the model has simply memorised the exam. The fix is to keep fresh, held-out tests and to evaluate on tasks the model has never seen — which is harder than it sounds when models read the whole internet.
Red teaming
Red teaming is adversarial evaluation: humans (or other models) actively try to make the system fail — jailbreaks, prompt injection, harmful requests, edge cases, social engineering. It is the security mindset applied to a language model. A system is only as safe as the worst thing a motivated user can make it do.
Evals are the interface
For agents, evaluation gets harder: there is no single right answer, outcomes are stochastic, and success may take many steps. You need trajectory evals — did the agent take a sane path, not just reach a lucky end state?