Skip to content

LLM as a judge

Using a language model to score another model's output, what it is good at, and the three failure modes that make an unchecked judge worthless.

Give a model a conversation and a question about it, and take its answer as a score. It is the only way to score language at volume, and it is unreliable in specific, knowable ways.

What it is good at

Questions with a defensible answer in the transcript. Did the agent say it was not a person when asked. Did it quote a price. Did it answer the question that was asked.

Three ways it goes wrong

Self-preferenceA model scores text from its own family more generously.
Position and verbosity biasOrder and length move the score independently of quality.
Confident ambiguityGiven a vague question it answers firmly anyway, differently each time.

None of these is visible in the output. A wrong score looks exactly like a right one, which is what makes an unmeasured judge worse than no judge: it produces numbers people act on.

What makes one trustworthy

  • Several models from different families scoring independently.
  • A separate step to settle a disagreement, rather than an average.
  • Agreement with human reviewers measured on a reviewed sample and published.
  • Every score citing the turn it came from, so a reader can check it.
The thing that built the agent should not be the only thing scoring it. A judge inside the same stack, from the same model family, prompted by the same team, shares the blind spots it is supposed to find.
This page as markdown