In 2023, Zheng and colleagues measured something that is often cited to defend using a language model as a judge: a strong model agrees with people in over 80% of cases, the same level of agreement that exists between people. It is a real figure, and it is easy to stretch it further than it goes.
It gets stretched because it is read as "the judge is right 80% of the time". It measures something else: in what share of comparisons between two answers from a chat assistant the judge picked the same one as a human expert. That says nothing about your criterion, your cases or what happens when the judge is wrong on a decision somebody has to defend.
The thesis of this piece is that an automated judge that has not been calibrated against a person is an opinion dressed up as a metric. You can use it where the criterion is binary and calibrated. Where the decision carries risk, financial, legal or in front of a customer, a person signs, and the judge prepares the case.
What the 80% shows and what it does not
The study is worth reading with a magnifying glass. The agreement the authors cite was measured with GPT-4 as the judge on MT-bench, a set of open-ended questions, comparing two answers and using votes from human experts. Without ties, GPT-4 and the experts agreed on 85%; among the experts themselves, on 81%.
Three limits, which the authors acknowledge or which follow from the design. The study evaluates chat assistants and focuses on whether the answer is helpful; the authors note that it largely neglects safety, honesty and harmlessness. It folds several dimensions, such as accuracy, relevance and creativity, into a single score. And it compares pairs of answers: it does not decide whether an action by an agent can be executed. None of this invalidates the result. But none of the three looks like "can this agent approve the refund?".
The judge's biases
The same study documents that an automated judge has biases of its own. The table lists the ones it names, with what the sources say is done about each.
| Bias | What the judge does | What is done |
|---|---|---|
| Position | Favours an answer because of the order it is shown in | Call it twice with the order swapped and declare a winner only if it wins in both orders (Zheng et al.) |
| Verbosity | Prefers the longer answer | Control answer length (OpenAI); use a reference answer and step-by-step reasoning, with limited effectiveness in some cases (Zheng et al.) |
| Self-enhancement | May favour answers from its own model | The authors could not determine whether it appears in their study. A precaution added by onext, not from the sources: the judge should not be the same model that wrote the answer |
| Limited reasoning | Fails on tasks that require reasoning in order to score | Whenever possible, a criterion checked by code rather than by the judge |
onext's own work, based on Zheng et al. (2023) and OpenAI's evaluation best practices guide; the self-enhancement row includes a precaution added by onext
The last row is the most practical. If a criterion can be checked with code, checking it with code is cheaper, more reproducible and easier to debug than asking a model. The judge keeps what code cannot score. In evals on every pull request (automated tests that score what an agent does) we explain where each type of grader fits in continuous integration (CI).
Why raw agreement misleads
This is the mistake that is hardest to see. Hamel Husain, in his guide to judges, puts it plainly: raw agreement can be misleading when classes are imbalanced, and he recommends treating the human labels as ground truth and reporting the judge's true positive rate and true negative rate separately.
That is why calibration is measured with two numbers: how many of the cases the person marked as good the judge also marks as good, and how many of those the person marked as bad the judge also marks as bad. The second one is usually what matters, because failures are what reach the customer.
How to calibrate a judge
Husain's guide proposes a sequence that can be followed without special tooling. OpenAI's documentation agrees on the essentials: validate agreement against your human labels before optimising for cost or latency, and scale up only when the judge is faster, cheaper and consistently agrees with human annotations. Anthropic's guide to evals adds the underlying reason: model-based graders are non-deterministic, cost more than code and need calibrating with human graders.
- One reference person. Husain asks you to identify a principal domain expert, whose judgment defines what an acceptable result is. Not a committee: someone.
- Binary labels with a critique. The person marks each case as pass or fail and writes why, in a sentence or two. Those critiques are the best examples for the judge.
- One judge per dimension. Anthropic recommends scoring each dimension with an isolated judge instead of asking a single one to score them all.
- An "Unknown" way out. Anthropic suggests giving the judge a way out, for example returning "Unknown" when it does not have enough information, to reduce hallucinations. Those cases go to a person.
- Measure against the labels. True positives and true negatives separately, with the threshold the company accepts written down before looking at the result.
With that you can say something the study's 80% does not allow: "for this criterion and these cases, the judge agrees with the reference person this well, and fails on these kinds of case". How to build the set of cases you measure against is in the golden set, and what "better" means when you compare two versions, in rubric and baseline.
Where a person signs
Calibrating a judge does not make it accountable. The sources ask for human review, but none says where to draw the line: it is a governance decision. This is a proposed split, by type of decision and not by how good the judge is.
| Type of decision | Example | What the judge does | Who decides |
|---|---|---|---|
| Binary, checkable criterion and a calibrated judge | The answer cites the source; the format is the one requested | Scores and blocks | The judge, with a sample a person reviews from time to time |
| Criterion with nuance | Usefulness, tone, clarity of an explanation | Flags and informs; does not block | A person looks at the trend and the doubtful cases |
| Decision with risk | Money, personal data, a legal obligation or a message to the customer | Prepares the case and orders it; does not resolve | A named person, and it is recorded |
onext's own work, based on Anthropic's and OpenAI's recommendation to calibrate the judge and keep human review; the split by type of decision is a proposal, not a rule from the sources
The line moves over time, and only in a direction you can demonstrate: a criterion with nuance can become binary once it has been rewritten until two people agree. A risk criterion does not become automatic because the judge has been right a lot; it does so, if it does, through a written decision by whoever answers for it. Where to sign when agents write the code is in the specification is where you sign off, and how a specification's criteria become the eval that stops a merge, in your spec is already an eval.
Risk A calibrated judge drifts. Anthropic notes that, once validated, occasional human review is enough, but insists on continuing to read transcripts. Calibration is a periodic habit, not a milestone. And watching tokens and latency is not measuring quality, as we develop in the quality gap in agents in production.
Where to start this week
If your agent today has a judge nobody has checked against a person, one person on the team can start like this:
- Pick one criterion. One, binary, that a judge scores today. Not the most important one: the clearest one.
- Label a sample. The reference person marks between 20 and 50 real cases as pass or fail, with a critique. Include the failures that already slipped through.
- Measure the judge against it. True positives and true negatives separately. Note which kind of case it fails on.
- Write the signing rule. Which decisions the judge may score and block, which it only informs and which a named person resolves.
- Set the periodic review. A sample every week or at every model change, and transcripts read by a person.
The range of 20 to 50 cases does not come from any study on judges: it is the order of magnitude Anthropic's guide says an eval suite can start with. It is a good start, not a statistical guarantee; with few cases, the result reads as a hint, not as a figure.
Frequently asked questions
What is LLM-as-a-judge?
It is using a language model to score the answers of another model or of an agent. According to OpenAI's documentation, it can be done by comparing two answers, scoring a single one or scoring it against a reference answer. It is cheaper and scales better than human evaluation, but it has to be validated against people.
Can you trust an LLM as a judge?
It depends on what you ask it and on whether you have measured it. In the study by Zheng and colleagues, GPT-4 agreed with human experts on 85% of the no-tie comparisons in MT-bench, against 81% among the people themselves. That is a good result for that specific case, not for any business criterion. OpenAI's documentation sums it up: no strategy is perfect and the quality of the judge varies with the problem.
Which biases does an automated judge have?
The study by Zheng and colleagues names three: position (it favours an answer because of the order it is shown in), verbosity (it prefers longer ones) and self-enhancement, although its authors say they could not determine whether the last one appears in their study. They also point to limited reasoning ability. OpenAI mentions the first two.
How do you calibrate a judge?
A person with domain expertise labels cases as pass or fail and explains why. The judge is adjusted until it matches those labels, and it is measured with the true positive rate and the true negative rate separately, not just with raw agreement. Hamel Husain warns that raw agreement misleads when classes are imbalanced.
Can an automated judge approve a decision with risk?
The sources neither forbid nor allow it: what they ask for is calibration and continued human review. Where to draw the line is a governance decision for each company. A reasonable rule is that, when a decision involves money, personal data, legal obligations or the customer, the judge prepares and orders the case, and a named person signs.
Is it better to score from 1 to 5 or with pass/fail?
The sources lean towards binary. Husain advises against 1-to-5 scales because they are not actionable, and OpenAI recommends pairwise comparison or pass/fail as more reliable. A binary criterion is also the one that can be calibrated clearly against a person.
Sources
- Lianmin Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena", arXiv 2306.05685, 9 June 2023 (revised 24 December 2023).
- Hamel Husain, "Using LLM-as-a-Judge For Evaluation: A Complete Guide", 29 October 2024 (modified 1 September 2026).
- Mikaela Grace, Jeremy Hadfield, Rodrigo Olivares and Jiri De Jonghe (Anthropic), "Demystifying evals for AI agents", 9 January 2026.
- OpenAI, "Evaluation best practices".

Bernat López is founder and CEO of onext, an AI boutique. He helps development and product teams work with AI with method —specification before building, a person who decides where there is risk and Spec-Driven Development— and applies to his own company what he proposes: onext runs on its own agentic system.
LinkedIn →