LLM as a judge
Using a language model to score another model's output, what it is good at, and the three failure modes that make an unchecked judge worthless.
Give a model a conversation and a question about it, and take its answer as a score. It is the only way to score language at volume, and it is unreliable in specific, knowable ways.
What it is good at
Questions with a defensible answer in the transcript. Did the agent say it was not a person when asked. Did it quote a price. Did it answer the question that was asked.
Three ways it goes wrong
| Self-preference | A model scores text from its own family more generously. |
|---|---|
| Position and verbosity bias | Order and length move the score independently of quality. |
| Confident ambiguity | Given a vague question it answers firmly anyway, differently each time. |
None of these is visible in the output. A wrong score looks exactly like a right one, which is what makes an unmeasured judge worse than no judge: it produces numbers people act on.
What makes one trustworthy
- Several models from different families scoring independently.
- A separate step to settle a disagreement, rather than an average.
- Agreement with human reviewers measured on a reviewed sample and published.
- Every score citing the turn it came from, so a reader can check it.
The thing that built the agent should not be the only thing scoring it. A judge inside the same stack, from the same model family, prompted by the same team, shares the blind spots it is supposed to find.
This page as markdown