Skip to content

This is a community translation of the original Chinese text. The translation may contain inaccuracies. When in doubt, please refer to the original Chinese version.

Evals Are the Probation Period

An agent gets examined twice, in two different stages. Collapsing the two exams into one is the most common confusion in evals.

The first exam comes before hiring: the general leaderboards, the equivalent of a candidate's school tests and standardized exams. It measures general ability and decides whether the candidate deserves a degree — nothing to do with any specific job. This exam matters, but once it's over, the candidate has graduated. They take the transcript job-hunting, and a transcript is not a work permit.

The second exam comes after hiring, and it's the subject of this piece: the probation period. Human companies set probation periods precisely because degrees and rankings say nothing about job competence. Probation doesn't look at class rank; it puts the person in the real job, doing real work, and asks whether they measure up. Evals should be this second exam for agents — not a rereading of the first exam's transcript.

Probation tests the job, not the rank

Getting the second exam right means designing it around the role: test the agent on this job's real tasks, real data, and real bar for passing — not on some generic leaderboard that has nothing to do with the role. The eval for an invoice-extraction role should measure extraction accuracy on real invoices and whether the agent escalates when it should. It should not measure its general reasoning score.

An eval detached from the job cannot measure competence

From this follows a rule: an eval detached from the job can't tell you whether an agent is good. It can only tell you the agent's general tier.

General tier is one input when picking a model, but it is not a verdict on job competence. Treating a high score on a general eval as "this agent can handle my role" is using the pre-probation resume in place of the probation itself. What you need to build is an assessment grown out of the job — one that tests how well this work gets done, not where the candidate ranks on generic problems.