Agreement with people
How we measure whether our scores match what your reviewers would have said.
A number you cannot audit is not evidence. So the scoring is measured against people: your reviewers score conversations by hand, and we compare what the judge said with what they said, rule by rule and conversation by conversation.
Where the human answers come from
From your own reviewers, inside the product. A reviewer opens a conversation, reads it, and records what they think the answer was. The product builds them a queue of conversations worth reading, so the examples are not only the easy ones.
Those hand-marked examples are the only thing the figures below are measured against. We do not buy a set in, and we never measure the judge against itself.
What one comparison is
One rule, on one conversation, with an answer from the judge and an answer from a person. Agreement is counted per score area over a period, and the number of examples behind it is published beside it, because a figure from nine examples and a figure from nine hundred are not the same claim.
The four figures
| Matches people | How often the judge's verdict matches what a person would have said, measured against conversations a person hand-scored. |
|---|---|
| Flag accuracy | When the judge flags a failure, how often it's a real one. Higher is better: a low number means a lot of false alarms. |
| Catch rate | Of the real failures, how many the judge catches. Higher is better: a low number means a lot of missed problems. |
| Score movement | How much a score moves on its own when nothing about the conversation changed. A change smaller than this is not real movement. |
The first three climb toward a perfect 1.00. The last one is movement, so a smaller number is the better one.
What we leave out, and say that we left out
Some answers a reviewer gives are about a rule that needed no AI at all. Those have a computable right answer, so they say nothing about how good the judge is, and they are held out of every figure above. How many were held out is published anyway, because an exclusion nobody can see cannot be told apart from choosing the numbers.
When a reviewer contradicts one of those, it is not a disagreement about judgement. It means we read something out of the platform wrongly, and it is filed against us as a defect. That count is published too.
Too few examples is not a pass
Until a score area has enough hand-marked examples, its figure reads as still settling rather than as good. High agreement over a handful of examples is not evidence of anything, and a workspace that has marked nothing by hand yet is told exactly that instead of being shown a flattering number.
Read next
- What gets scored · Why the fast rules run on every conversation and the ones that cost money run on a share.
- What a score is · What a number here means, what it does not mean, and what you must not use it for.
- How it works · The six stages a conversation passes through, from arriving to being scored.