AI Foundationspredict · compress · act
act VIII

The Alignment Layer

Benchmarks are how the field measures progress and how it fools itself. Good evaluation is a discipline, not a leaderboard.

32

How Do We Even Know?

before this →Getting the Goal Right

You cannot improve what you cannot measure, and in AI it is dangerously easy to measure the wrong thing. Evaluation is the craft of finding out what a model can really do.

capability evalscan it do the task? accuracy, pass@k, win ratesafety evalsdoes it refuse harm, resist jailbreaks, stay in scope?deployment evalsdoes it work in your domain, with your data, at your cost?
the evaluation stack: capability, safety, and the gap between them

Benchmarks and their lies

A benchmark is a fixed test everyone runs. It is useful and it is inevitably gamed. Once a benchmark becomes a target, models are trained on data that resembles it, and the score stops predicting real ability. Contamination is the acute version: test questions leak into the training corpus, and the model has simply memorised the exam. The fix is to keep fresh, held-out tests and to evaluate on tasks the model has never seen — which is harder than it sounds when models read the whole internet.

Red teaming

Red teaming is adversarial evaluation: humans (or other models) actively try to make the system fail — jailbreaks, prompt injection, harmful requests, edge cases, social engineering. It is the security mindset applied to a language model. A system is only as safe as the worst thing a motivated user can make it do.

Evals are the interface

between research and the real world

For agents, evaluation gets harder: there is no single right answer, outcomes are stochastic, and success may take many steps. You need trajectory evals — did the agent take a sane path, not just reach a lucky end state?

the disciplineA model’s headline score tells you almost nothing about whether it is right for your task. Build your own eval before you build your product.
DEFINEwhat does success mean here?
MEASUREheld-out, uncontaminated, repeated
ADVERSARYtry to break it on purpose
introduces →benchmarkevaluationcontaminationred teaming
← previousGetting the Goal Rightnext →From Answer to Action