# Ground truth

> The human judgement an automated score is measured against, and why calling it truth hides the most important question about it.

The set of human decisions an automated check is compared against. It is what agreement is agreement with.

## The word is doing too much work

It is not truth. It is what some people decided, on some conversations, on some day, under some reading of a rule.

Two careful reviewers reading the same transcript against the same rule will disagree on a meaningful share of cases, and neither is making a mistake. The rule was ambiguous.

## The questions that matter about it

- Who labelled it, and did they know the product or the policy.
- How much they disagreed with each other before anything was settled.
- How disagreements were resolved, and whether resolution was recorded.
- Whether the labelled sample resembles production, or only the easy part of it.

> If your own reviewers agree with each other 85 percent of the time, no automated check can be shown to agree with them much above 85 percent. The ceiling on a judge is the consistency of the people it is measured against.

That ceiling is a reason to measure human disagreement, not a reason to hide it. [Label noise](/glossary/label-noise)

---

Source: https://evidova.com/glossary/ground-truth
