# LLM as a judge

> Using a language model to score another model's output, what it is good at, and the three failure modes that make an unchecked judge worthless.

Give a model a conversation and a question about it, and take its answer as a score. It is the only way to score language at volume, and it is unreliable in specific, knowable ways.

## What it is good at

Questions with a defensible answer in the transcript. Did the agent say it was not a person when asked. Did it quote a price. Did it answer the question that was asked.

## Three ways it goes wrong

| | |
|---|---|
| Self-preference | A model scores text from its own family more generously. |
| Position and verbosity bias | Order and length move the score independently of quality. |
| Confident ambiguity | Given a vague question it answers firmly anyway, differently each time. |

None of these is visible in the output. A wrong score looks exactly like a right one, which is what makes an unmeasured judge worse than no judge: it produces numbers people act on.

## What makes one trustworthy

- Several models from different families scoring independently.
- A separate step to settle a disagreement, rather than an average.
- Agreement with human reviewers measured on a reviewed sample and published.
- Every score citing the turn it came from, so a reader can check it.

> The thing that built the agent should not be the only thing scoring it. A judge inside the same stack, from the same model family, prompted by the same team, shares the blind spots it is supposed to find.

---

Source: https://evidova.com/glossary/llm-as-a-judge
