Judge agreement
How far an automated judge matches the humans it stands in for, which statistic to quote, and what a number without a method is worth.
If a model scores your conversations instead of a person, the only question that matters is how often it reaches the same answer the person would have.
That number is measurable. It is measured by having people score a sample, then comparing. There is no way to know it without doing that.
It is not a reliability question
Agreement coefficients were built for two interchangeable raters, neither of whom is the truth. That is not this situation.
You are asking whether to deploy the judge, with human labels as the reference. That is a question about accuracy, so sensitivity and specificity carry more than a reliability coefficient does.
Sensitivity and specificity also transfer across workspaces with different failure rates. Precision does not, so derive it at your own prevalence rather than quoting someone else's.
What to quote
- Prevalence, so the reader knows how rare the thing being caught is.
- Raw agreement, because everyone can read it.
- A chance-corrected coefficient, because raw agreement alone flatters.
- Recall on the failure class, because missing real failures is the expensive error.
- The sample size, because a coefficient from 30 labels means almost nothing.
A number without a method
Published agreement figures in this category are common. Published methods are not. A percentage with no dataset, no sample size and no definition of agreement cannot be checked or compared.
Ask what counted as agreement, how many items, who labelled them, how disagreements were settled, and whether the figure is for one customer or pooled across many.
Hold us to this too. Our figures are per workspace, with the sample size beside them. Agreement with people
This page as markdown