In its guide to evals for agents, published in January 2026, Anthropic's engineering team writes that defining eval tasks is one of the best ways to stress-test whether a product's requirements are concrete enough to start building. They are thinking of people who build agents. It applies just as well to a team whose agents write the code.
In Spec-Driven Development (SDD), the specification is the source of truth and the code is derived from it. In practice, though, many teams keep two artefacts that never talk to each other: the specification, written at the start and read once, and the tests, written afterwards, often generated by the agent itself and often derived from the code rather than from the specification.
The argument of this piece is that those two artefacts should be one. If every acceptance criterion is written so that a machine or a specific person can score it without interpretation, the specification is already the eval: the set of checks that decides whether a change reaches the main branch. And the uncomfortable corollary: without a criterion that can be scored, the specification is not finished, however well written it is.
Two documents that should be one
The idea is not new; what is new is that it is now cheap. The method document of spec-kit, GitHub's toolkit for SDD, sums it up in one line: "Acceptance scenarios become tests." It adds that test scenarios are not written after the code; they are part of the specification that generates both the implementation and the tests.
Whether that actually happens depends on how the criterion is written. Anthropic's guide defines a task as a single test with defined inputs and success criteria, and a grader as the logic that scores some aspect of the result. An acceptance criterion with concrete inputs and an expected result already is an evaluation task. What it lacks to be an eval is the grader: who or what says "passes" or "fails".
When that grader does not come from the criterion, it comes from the code. And a test written from the code does not test it, it describes it: it confirms what the code does, including what it does wrong. We explain this in detail in AI-generated tests. The specification as an eval is the other side of that problem: the criterion exists before the code and is the only thing that can judge it.
The two-reader test
How do you know whether a criterion can be scored? Anthropic's guide gives a simple yardstick for a good evaluation task: two domain experts would independently reach the same pass/fail verdict. And it warns what happens otherwise: ambiguity in task specifications becomes noise in metrics.
The two-reader test: if two people who know the product can read a criterion, look at the result and disagree on whether it is met, that criterion is not a criterion yet. It is an intention.
Anthropic's documentation on success criteria makes the same point in other words: specific and measurable. Its example of a bad criterion is "the model should classify sentiments well"; the good one sets the metric, the threshold and the dataset it is measured against. In a development team the translation is direct. These are examples made up to illustrate the point, not from a real project:
| Criterion that cannot be scored | The same criterion, ready to be an eval |
|---|---|
| "Invoice export must be fast" | Given a file of 10,000 invoices, when it is exported in the continuous integration environment, then the export finishes in under 5 seconds |
| "Errors must be clear" | Given an invoice without a tax ID, when it is sent to the API, then it responds 422 and the tax_id field appears in the error list |
| "It must respect permissions" | Given a user without the billing role, when they request the export, then they get 403 and no file is generated |
Illustrative examples, onext's own elaboration. The "given, when, then" format is BDD's GIVEN/WHEN/THEN
The format of the right-hand column is no accident. Birgitta Böckeler of Thoughtworks, analysing SDD tools, describes how one of them, Kiro, structures requirements as user stories with acceptance criteria in GIVEN… WHEN… THEN… format. That format forces you to state the input, the action and the result, which is exactly what a grader needs.
Three kinds of criterion, three graders
Not every criterion is scored the same way, and forcing everything into an automated test is the opposite mistake to having none. Anthropic's guide lists the strengths of code-based graders —fast, cheap, objective, reproducible— and also their limits: they are brittle to valid variations that don't match expected patterns exactly, and they lack nuance. For what has nuance it proposes models as judges, frequently calibrated against expert human judgment. Applied to a specification:
| Kind of criterion | Who scores it | What it does at merge time |
|---|---|---|
| Deterministic | Code: a unit, integration or contract test derived from the criterion | Blocks. If it fails, the change does not get in |
| With nuance | A rubric written into the specification; if a model scores it, calibrated against people first | Informs. It only blocks once the judge is calibrated and the threshold agreed |
| Risk or business | A named person, set out in the specification | Signs off. Without their sign-off, it does not get in |
onext's own elaboration, based on Anthropic's classification of graders (code, model and human)
The third row is the one most often forgotten. Some criteria should not be decided by any automated grader: those that touch money, personal data, permissions or a debatable business rule. There the eval is a person, and the specification has to say who. It is the idea we develop in the specification is where you sign off: you don't sign off everything, because approving everything isn't control, it's a bottleneck; you sign off where there is risk.
The second row has its own method. A rubric without a set of reference cases is an opinion dressed as a metric; how to build that set and what "works better" means is covered in the golden set and in rubric and baseline.
The merge gate: new criteria fail first, earlier ones keep passing
With the criteria classified, the merge gate is built on two rules, and neither is ours.
New criteria fail first. The spec-kit method demands strict TDD: no implementation code is written until the tests are written, a person has approved them and they have been confirmed to fail. A test that passes before the code exists is not checking anything. Applied to the specification as an eval: the tests that come out of the new criteria are written before the code (an agent can draft them), and a person reviews them against the criterion, not against the implementation, and they must fail red before work starts.
Earlier criteria keep passing. The criteria from specifications already merged form the regression suite. Anthropic's guide says a regression suite should have a nearly 100% pass rate, and that capability evals that are passed comfortably can "graduate" to regression. According to the same guide, it is the structure used to score coding agents in SWE-bench Verified: a solution passes only if it fixes the failing tests without breaking existing ones.
With that, the question you ask when reviewing a pull request changes. It is no longer "am I convinced by this code?", which becomes the bottleneck when an agent produces more changes per hour, but "are the specification's criteria covered and green, and have the people who had to sign off done so?". Human review doesn't disappear; it concentrates where it adds value. We develop this in the review that still expects an author.
What this doesn't solve
The limits deserve the same clarity as the argument. Böckeler points out the first: even with all the files, templates, workflows and checklists of SDD tools, she frequently saw the agent ultimately not follow all the instructions. A longer specification does not guarantee compliance. That is precisely why control doesn't lie in the length of the specification, but in checking every criterion before merging.
The second limit is the flip side of the first: an eval only passes what it measures. Whatever the specification leaves out, no test will catch, and a green dashboard can give the same false sense of control that Böckeler raises about SDD tools themselves. The two-reader test helps make sure each criterion is well written; it doesn't tell you whether criteria are missing. That is still the job of whoever knows the product.
And the third is cost: writing criteria that can be scored takes longer up front than writing intentions. We don't have a measured figure for how much, and we are not going to make one up. What can be reasoned is where you pay if you don't: in a review that argues about what the specification meant once the code is already written.
If you want to know where this starts in your team, take the last specification you signed off as good and go through its criteria one by one with a single question: who or what scores it. The ones with an answer are already your eval. The ones without are the pending work, and they tend to be the same ones that cause the longest arguments in review. It fits the processes perspective of the four perspectives of a development team working with AI: ready is the specification and done is the verification against it. And if the whole method is new to you, start with what Spec-Driven Development is.
The question for your next team meeting: how many criteria in your last specification could be scored by someone who didn't write it?
Frequently asked questions
What does it mean for a specification to be an eval?
It means every acceptance criterion in the specification is written with concrete inputs and a result that can be scored, so the check that decides whether a change gets in comes straight out of it. The specification stops being documentation you read before coding and becomes what runs before you merge.
How do I know whether an acceptance criterion is verifiable?
With the two-reader test: if two people who know the product can read the criterion, look at the result and independently reach different verdicts, the criterion is not finished. Anthropic's guide to evals uses the same yardstick for a good evaluation task and warns that ambiguity in the specification becomes noise in the metrics.
Do all acceptance criteria have to become automated tests?
No. Deterministic criteria are checked with code and can block the merge. Criteria with nuance, such as how clear a message is, are scored with a rubric and, if a model acts as the judge, calibrated against people. And criteria that depend on a business judgement or carry risk are signed off by a named person. Forcing everything into a test produces brittle tests that fail on valid variations.
Who should write the test that comes out of an acceptance criterion?
An agent can write it, but before the code and from the criterion, never from code that has already been generated. A person reviews that test against the criterion, not against the implementation. A test derived from the code only describes what the code does, including what it does wrong.
Which evals should block a merge?
Two groups. The new criteria in the specification, which must fail before implementation and pass afterwards. And the criteria from earlier specifications, which form the regression suite: according to Anthropic's guide to evals, a regression suite should have a nearly 100% pass rate. It is the same structure SWE-bench Verified uses to accept a solution.
Is a more detailed specification enough to make the agent follow it?
No. Birgitta Böckeler of Thoughtworks reports that even with templates, checklists and complete workflows she frequently saw the agent ultimately not follow all the instructions. More specification does not guarantee compliance; what gives you control is checking every criterion before merging.
Sources
- Mikaela Grace, Jeremy Hadfield, Rodrigo Olivares and Jiri De Jonghe (Anthropic), "Demystifying evals for AI agents", 9 January 2026.
- Anthropic, Claude documentation, "Define success criteria and build evaluations".
- GitHub, spec-kit, "Specification-Driven Development (SDD)".
- Birgitta Böckeler (Thoughtworks), "Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl", martinfowler.com, 15 October 2025.

Bernat López is founder and CEO of onext, an AI boutique. He helps development and product teams work with AI with method —specification before building, a person who decides where there is risk and Spec-Driven Development— and applies to his own company what he proposes: onext runs on its own agentic system.
LinkedIn →