Skip to content

Label noise

Disagreement and error inside your own human labels, which puts a hard ceiling on the agreement any automated judge can be measured at.

Human labels contain mistakes, fatigue, drift and honest disagreement. That is label noise, and it is present in every reviewed set anyone has ever built.

The ceiling it sets

If your reviewers agree with each other 85 percent of the time, an automated check cannot be shown to agree with them much beyond 85 percent, however good it is.

A system reporting 97 percent agreement against labels that are 85 percent self-consistent is reporting something other than what it claims.

Where it comes from

Ambiguous rulesThe largest source, and the most fixable. Two readings, both defensible.
Reviewer driftThe same person labels differently in week four than week one.
FatigueAccuracy falls across a long session, usually toward the easy answer.
Missing contextA reviewer who cannot see the audio or the account scores a different thing.

What to do with it

Measure it before you measure anything else. Have two people label the same sample and compute their agreement. That number is the ceiling for everything downstream.

Then treat high disagreement on a rule as a signal the rule needs rewording, rather than as a reviewer problem.

This page as markdown