Skip to main content
onext technology
AI 7 October 2026 - 9 min read

LLM-as-judge: when to trust the judge and where a person has to sign

An automated judge that has not been calibrated against a person is an opinion dressed up as a metric. It works where the criterion is binary and calibrated; for decisions with risk, a person signs.

Bernat López
Founder and CEO of onext
Diagram of an automated judge: a grid of cases enters a navy bar; one lane leaves with automated checks and the other with a ticked box next to a person signing, in blue and navy on white

In 2023, Zheng and colleagues measured something that is often cited to defend using a language model as a judge: a strong model agrees with people in over 80% of cases, the same level of agreement that exists between people. It is a real figure, and it is easy to stretch it further than it goes.

It gets stretched because it is read as "the judge is right 80% of the time". It measures something else: in what share of comparisons between two answers from a chat assistant the judge picked the same one as a human expert. That says nothing about your criterion, your cases or what happens when the judge is wrong on a decision somebody has to defend.

The thesis of this piece is that an automated judge that has not been calibrated against a person is an opinion dressed up as a metric. You can use it where the criterion is binary and calibrated. Where the decision carries risk, financial, legal or in front of a customer, a person signs, and the judge prepares the case.

What the 80% shows and what it does not

The study is worth reading with a magnifying glass. The agreement the authors cite was measured with GPT-4 as the judge on MT-bench, a set of open-ended questions, comparing two answers and using votes from human experts. Without ties, GPT-4 and the experts agreed on 85%; among the experts themselves, on 81%.

85% agreement between GPT-4 and human experts on MT-bench, without ties (Zheng et al., 2023)
81% agreement among the human experts themselves on the same set

Three limits, which the authors acknowledge or which follow from the design. The study evaluates chat assistants and focuses on whether the answer is helpful; the authors note that it largely neglects safety, honesty and harmlessness. It folds several dimensions, such as accuracy, relevance and creativity, into a single score. And it compares pairs of answers: it does not decide whether an action by an agent can be executed. None of this invalidates the result. But none of the three looks like "can this agent approve the refund?".

The judge's biases

The same study documents that an automated judge has biases of its own. The table lists the ones it names, with what the sources say is done about each.

Bias What the judge does What is done
Position Favours an answer because of the order it is shown in Call it twice with the order swapped and declare a winner only if it wins in both orders (Zheng et al.)
Verbosity Prefers the longer answer Control answer length (OpenAI); use a reference answer and step-by-step reasoning, with limited effectiveness in some cases (Zheng et al.)
Self-enhancement May favour answers from its own model The authors could not determine whether it appears in their study. A precaution added by onext, not from the sources: the judge should not be the same model that wrote the answer
Limited reasoning Fails on tasks that require reasoning in order to score Whenever possible, a criterion checked by code rather than by the judge

onext's own work, based on Zheng et al. (2023) and OpenAI's evaluation best practices guide; the self-enhancement row includes a precaution added by onext

The last row is the most practical. If a criterion can be checked with code, checking it with code is cheaper, more reproducible and easier to debug than asking a model. The judge keeps what code cannot score. In evals on every pull request (automated tests that score what an agent does) we explain where each type of grader fits in continuous integration (CI).

Why raw agreement misleads

This is the mistake that is hardest to see. Hamel Husain, in his guide to judges, puts it plainly: raw agreement can be misleading when classes are imbalanced, and he recommends treating the human labels as ground truth and reporting the judge's true positive rate and true negative rate separately.

The sum (our own arithmetic) If 95 out of 100 answers from your agent are correct, a judge that always says "correct" agrees with the person in 95% of cases and does not catch a single failure. 95% agreement sounds calibrated. With that figure alone, you do not know whether the judge sees the errors.

That is why calibration is measured with two numbers: how many of the cases the person marked as good the judge also marks as good, and how many of those the person marked as bad the judge also marks as bad. The second one is usually what matters, because failures are what reach the customer.

How to calibrate a judge

Husain's guide proposes a sequence that can be followed without special tooling. OpenAI's documentation agrees on the essentials: validate agreement against your human labels before optimising for cost or latency, and scale up only when the judge is faster, cheaper and consistently agrees with human annotations. Anthropic's guide to evals adds the underlying reason: model-based graders are non-deterministic, cost more than code and need calibrating with human graders.

  1. One reference person. Husain asks you to identify a principal domain expert, whose judgment defines what an acceptable result is. Not a committee: someone.
  2. Binary labels with a critique. The person marks each case as pass or fail and writes why, in a sentence or two. Those critiques are the best examples for the judge.
  3. One judge per dimension. Anthropic recommends scoring each dimension with an isolated judge instead of asking a single one to score them all.
  4. An "Unknown" way out. Anthropic suggests giving the judge a way out, for example returning "Unknown" when it does not have enough information, to reduce hallucinations. Those cases go to a person.
  5. Measure against the labels. True positives and true negatives separately, with the threshold the company accepts written down before looking at the result.

With that you can say something the study's 80% does not allow: "for this criterion and these cases, the judge agrees with the reference person this well, and fails on these kinds of case". How to build the set of cases you measure against is in the golden set, and what "better" means when you compare two versions, in rubric and baseline.

Where a person signs

Calibrating a judge does not make it accountable. The sources ask for human review, but none says where to draw the line: it is a governance decision. This is a proposed split, by type of decision and not by how good the judge is.

Type of decision Example What the judge does Who decides
Binary, checkable criterion and a calibrated judge The answer cites the source; the format is the one requested Scores and blocks The judge, with a sample a person reviews from time to time
Criterion with nuance Usefulness, tone, clarity of an explanation Flags and informs; does not block A person looks at the trend and the doubtful cases
Decision with risk Money, personal data, a legal obligation or a message to the customer Prepares the case and orders it; does not resolve A named person, and it is recorded

onext's own work, based on Anthropic's and OpenAI's recommendation to calibrate the judge and keep human review; the split by type of decision is a proposal, not a rule from the sources

The line moves over time, and only in a direction you can demonstrate: a criterion with nuance can become binary once it has been rewritten until two people agree. A risk criterion does not become automatic because the judge has been right a lot; it does so, if it does, through a written decision by whoever answers for it. Where to sign when agents write the code is in the specification is where you sign off, and how a specification's criteria become the eval that stops a merge, in your spec is already an eval.

Risk A calibrated judge drifts. Anthropic notes that, once validated, occasional human review is enough, but insists on continuing to read transcripts. Calibration is a periodic habit, not a milestone. And watching tokens and latency is not measuring quality, as we develop in the quality gap in agents in production.

Where to start this week

If your agent today has a judge nobody has checked against a person, one person on the team can start like this:

  1. Pick one criterion. One, binary, that a judge scores today. Not the most important one: the clearest one.
  2. Label a sample. The reference person marks between 20 and 50 real cases as pass or fail, with a critique. Include the failures that already slipped through.
  3. Measure the judge against it. True positives and true negatives separately. Note which kind of case it fails on.
  4. Write the signing rule. Which decisions the judge may score and block, which it only informs and which a named person resolves.
  5. Set the periodic review. A sample every week or at every model change, and transcripts read by a person.

The range of 20 to 50 cases does not come from any study on judges: it is the order of magnitude Anthropic's guide says an eval suite can start with. It is a good start, not a statistical guarantee; with few cases, the result reads as a hint, not as a figure.

Frequently asked questions

What is LLM-as-a-judge?

It is using a language model to score the answers of another model or of an agent. According to OpenAI's documentation, it can be done by comparing two answers, scoring a single one or scoring it against a reference answer. It is cheaper and scales better than human evaluation, but it has to be validated against people.

Can you trust an LLM as a judge?

It depends on what you ask it and on whether you have measured it. In the study by Zheng and colleagues, GPT-4 agreed with human experts on 85% of the no-tie comparisons in MT-bench, against 81% among the people themselves. That is a good result for that specific case, not for any business criterion. OpenAI's documentation sums it up: no strategy is perfect and the quality of the judge varies with the problem.

Which biases does an automated judge have?

The study by Zheng and colleagues names three: position (it favours an answer because of the order it is shown in), verbosity (it prefers longer ones) and self-enhancement, although its authors say they could not determine whether the last one appears in their study. They also point to limited reasoning ability. OpenAI mentions the first two.

How do you calibrate a judge?

A person with domain expertise labels cases as pass or fail and explains why. The judge is adjusted until it matches those labels, and it is measured with the true positive rate and the true negative rate separately, not just with raw agreement. Hamel Husain warns that raw agreement misleads when classes are imbalanced.

Can an automated judge approve a decision with risk?

The sources neither forbid nor allow it: what they ask for is calibration and continued human review. Where to draw the line is a governance decision for each company. A reasonable rule is that, when a decision involves money, personal data, legal obligations or the customer, the judge prepares and orders the case, and a named person signs.

Is it better to score from 1 to 5 or with pass/fail?

The sources lean towards binary. Husain advises against 1-to-5 scales because they are not actionable, and OpenAI recommends pairwise comparison or pass/fail as more reliable. A binary criterion is also the one that can be calibrated clearly against a person.

Sources

Bernat López
Written by
Bernat López
Founder and CEO of onext

Bernat López is founder and CEO of onext, an AI boutique. He helps development and product teams work with AI with method —specification before building, a person who decides where there is risk and Spec-Driven Development— and applies to his own company what he proposes: onext runs on its own agentic system.

LinkedIn →

Who signs today what your agent decides?

In Enterprise AI, in processes with risk, the assistant proposes and a person resolves from a queue of pending tasks. An audit record is kept.

See how we work