Skip to main content
onext technology
AI 6 October 2026 - 9 min read

Evals on every pull request: what runs per PR, what runs nightly and what it costs

"Run your evals on every change" is good advice until the pull request takes half an hour and the model bill soars. Not every eval should block a merge: regression, small and deterministic, runs on every PR; capability, with a judge and several trials, runs at night. And changing a prompt is changing code.

Bernat López
Founder and CEO of onext
Diagram of evals in CI: a short lane of fast checks up to the pull request gate and a nightly lane with repeated trials, in blue and navy on white

The guide to agent evals that Anthropic's engineering team published in January 2026 puts it plainly: automated evals are especially useful in CI/CD, running on each agent change and model upgrade as the first line of defence against quality problems. OpenAI's evaluation best practices say the same in other words: continuous evaluation, evals on every change and an eval set that grows over time.

The advice is right. The problem starts when a team applies it literally, with all its evals at once: the pull request takes as long as the slowest judge, the result changes from one run to the next with nobody having touched anything, and the token bill grows with every push. A few weeks later one of two things happens: someone marks the job as optional, or the team stops looking at the result. Either way, the evals still exist and no longer protect anything.

The thesis of this piece is that not every eval should block a pull request. Every PR runs the regression suite: small, deterministic, cheap and close to a 100% pass rate. At night, or before a model change, the capability evals run: with a model as judge, several trials per task and a cost that only makes sense to pay once a day. And one rule with no exceptions: a prompt, skill or model change is a code change, and it goes through the same gate.

Two different questions, two different suites

The distinction is not ours. Anthropic's guide separates two kinds of eval that answer different questions. Regression evals ask whether the agent still handles all the tasks it used to, and should have a pass rate close to 100%. Capability evals ask what the agent does well, and should start at a low pass rate, because they target the tasks it still struggles with.

That difference decides where each one runs. An eval that should always pass can block a merge: if it fails, something broke. An eval designed to fail often cannot block anything, because it would block almost every change. Mixing them in the same job is the quickest way to teach a team to ignore red.

Suite Question it asks Who grades it When it runs What it does at merge
Regression Does what used to work still work? Code: exact matches, tests, schemas, contracts On every pull request that touches what it evaluates Blocks. It should pass at close to 100%
Capability What does it do well, and is it improving? A rubric, often with a model as judge calibrated against people Nightly and before a model or version change Informs. You review the trend, not a single PR
Risk Can a machine decide this? A named person When the change touches money, personal data, permissions or a business rule Signs off. Without that signature, it does not go in

onext's own elaboration, based on the eval and grader types in Anthropic's guide (January 2026)

The third row is not an automated eval, and that is exactly why it belongs in the same table. It is the one we explain in your spec is already an eval: the specification's criteria say which grader applies to each thing, and some of them no machine can score.

What runs on every pull request

The suite that blocks a merge has to be fast, cheap and reliable, or the team will stop respecting it. Anthropic's guide lists the virtues of code-based graders: fast, cheap, objective, reproducible and easy to debug. That list is, literally, the specification of a pull request suite. In practice, four decisions keep it that way:

  • It only runs when it should. Promptfoo's CI/CD documentation (Promptfoo is an eval tool) shows a GitHub Actions workflow triggered on pull requests that change the prompts or the eval configuration, filtering by path. A stylesheet change has no reason to pay for an agent's evals.
  • What has not changed is not paid for again. The same example caches results under a key that depends on the content of the prompts. If the prompt is identical, the model call is not repeated.
  • The threshold is written down. Promptfoo offers two ways to break the build: fail on any eval error, or compute the pass rate and exit with an error if it falls below a target. For a regression suite, the sensible target is the one Anthropic gives: close to 100%.
  • Every trial starts clean. Anthropic insists that each trial be isolated and start from a clean environment. An eval that inherits state from the previous run fails or passes for reasons nobody will be able to reproduce.

If the pull request also runs an eval with a model as judge, its result is posted as a comment but does not block. It is the same rule we propose for nuanced criteria: they inform until the judge is calibrated and the threshold agreed. How to build the set of cases for that suite is covered in the golden set.

What waits for the night

Capability evals are expensive by design, not because of a flaw you can optimise away. Three reasons, all from Anthropic's guide:

  • They use judges that cost money. Model-based graders are non-deterministic, more expensive than code and need calibration with human graders to be accurate.
  • They need several trials per task. Because model outputs vary between runs, you run multiple trials to get more consistent results. Every extra trial is another full run of the agent.
  • They measure different things depending on how you count. pass@k measures the likelihood of at least one success in k attempts; pass^k, the probability that all k succeed. The more trials, the lower pass^k, because demanding consistency is a higher bar. If a customer is going to use the agent, the second one is what matters.

None of that fits in the minutes a developer is willing to wait in front of a pull request. It does fit in a nightly run, whose result you read in the morning as a trend: which tasks go up, which go down and which have been stuck for days. And it fits before a model change, which is exactly when Anthropic asks you to run the evals.

Graduation When a capability eval is passed comfortably, Anthropic suggests it "graduate" to the regression suite, which runs continuously to catch any drift. That way the pull request suite grows with what the agent has mastered, and the nightly suite keeps what it still struggles with.

Graduation has a limit the same guide points out: saturation. When an agent passes every solvable task in an eval, that eval leaves no room to measure improvement. The nightly suite needs new tasks, and the best ones come from real failures. What "better" means when you compare two versions is developed in rubric and baseline.

What it costs: the sum you have to do

We are not going to give a cost figure, because there is none that holds for every team and we have no measurement of our own that we can publish. What we can give is the sum. It is our own arithmetic, not from any source:

Cost of one run ≈ number of tasks × trials per task × (agent tokens + judge tokens) × price per token. If the run happens on every pull request, multiply as well by the day's pull requests that touch what is being evaluated.

The formula shows where the levers are. Trials per task multiply everything else, which is why they stay out of the pull request. The judge adds one call per response, which is why it only runs on the PR when it informs. Path filters and caching reduce the last factor, the number of runs, without touching the quality of the suite.

To have your own numbers instead of guesses, Anthropic suggests tracking latency, token usage, cost per task and error rates on a static bank of tasks. With two weeks of that data, deciding what goes into the pull request stops being an opinion. If your team builds products with AI, cost per useful task is also a business metric, as we explain in how to measure AI in a SaaS product.

≈ 100% pass rate a regression suite should have, according to Anthropic's guide to evals
20–50 simple tasks drawn from real failures: according to the same guide, a great start for a suite

A prompt is code, and it goes through the same gate

The rule most often skipped is not technical. In many teams, changing the agent's code goes through pull request, review and CI, but changing its prompt, a skill or the shared instructions is done directly, because "it's just text". And changing the model version is one line in a configuration file nobody reviews.

All four change the agent's behaviour as much as code does. Anthropic's guide asks for evals on every agent change and every model upgrade; Promptfoo's example triggers the workflow precisely when prompts change. If the team's skills and instructions live in the repository, as we argue in the practical guide to skills, the path filter covers them at no extra cost. If they live in a shared document outside the repository, nothing covers them.

There is a useful side effect: when a prompt change has to pass the regression suite, the team starts writing prompts more carefully, just as happened with code when tests arrived. Volume of cases matters more than the perfection of each one; Anthropic's documentation on success criteria puts it this way: more questions with slightly lower-signal automated grading are better than fewer questions hand-graded by people.

Risk No suite, however well split, catches everything. Anthropic uses the Swiss cheese model from safety engineering: no single evaluation layer catches every issue. CI evals are combined with production monitoring and with people who read transcripts; the guide insists that without reading them, you cannot know whether your graders are working.

Where to start this week

If you have no evals in CI today, or you have all of them in the same job, here is a sequence one person on the team can start:

  1. Inventory. List the evals that already exist, even loose scripts, and mark each one as regression or capability with Anthropic's question: should it always pass?
  2. Regression suite. Put together the regression evals that are graded by code and add the real failures from recent weeks, each one turned into a case. No judge and a single trial per task.
  3. Path trigger. Hook it to the pull request only when the agent's code, its prompts, its skills, its instructions or the model version change, with content-based caching and a written threshold.
  4. Nightly run. Move the capability evals there, with several trials per task and a calibrated judge, and note the trend every morning, not just the last result.
  5. Two weeks of measurement. Record latency, tokens and cost per task before expanding anything. With that data you decide what graduates and what stays nightly.

None of this requires a particular tool: the same logic works with Promptfoo, with another eval tool or with your own script in any CI system. What changes the outcome is the split, not the product. And if the tests that end up in the suite are generated by the agent, read first why a test written from the code does not test it.

Frequently asked questions

Should every eval run on every pull request?

No. Every pull request runs the regression suite: small, deterministic, cheap cases that check that what already worked still works, and that should pass at close to 100%. Capability evals, with a model as judge and several trials per task, are slower and more expensive, and run nightly or before a model change.

What is the difference between a regression eval and a capability eval?

According to Anthropic's guide to evals, a regression eval asks whether the agent still handles everything it used to, and should have a pass rate close to 100%. A capability eval asks what the agent does well, and should start at a low pass rate, because it targets what the agent still struggles with. When a capability eval is passed comfortably, it can graduate to the regression suite.

Can an eval with a model as judge block a merge?

Only if the judge has been calibrated against people and the threshold has been agreed. Anthropic notes that model-based graders are non-deterministic, more expensive than code and need calibration with human graders. Until then, their result informs the pull request but does not block it.

Should a prompt or model change go through CI?

Yes. A prompt, a skill, the team's shared instructions or the model version change behaviour just as code does. Anthropic's guide recommends running automated evals on every agent change and every model upgrade, as the first line of defence.

How much does it cost to run evals in CI?

There is no general figure: it depends on the number of tasks, the trials per task, the agent's and the judge's tokens and the model price. What you can do is measure it: Anthropic suggests tracking latency, token usage and cost per task on a static bank of tasks. With that, each team knows what it can afford on every pull request and what has to wait for the night.

How many cases do I need to start?

Not many. Anthropic's guide says that 20 to 50 simple tasks drawn from real failures are a great start. What matters is that every failure that slips into production becomes a new case in the suite.

Sources

Bernat López
Written by
Bernat López
Founder and CEO of onext

Bernat López is founder and CEO of onext, an AI boutique. He helps development and product teams work with AI with method —specification before building, a person who decides where there is risk and Spec-Driven Development— and applies to his own company what he proposes: onext runs on its own agentic system.

LinkedIn →

How does your team verify what the AI writes before production?

It fits the processes perspective of our free diagnostic, which looks at where your team loses speed or control today and what would change with a method of specification, AI implementation and verification. The method stays with your team.

See how we work

Without selling tools. Without stopping delivery.