Confidence interval
The range a measured rate could really sit in, how sample size moves it, and how many labels an agreement figure needs before it means anything.
Any rate measured on a sample is an estimate. The interval is the range the real value plausibly sits in, given how much you looked at.
Sample size is most of it
| 30 labels | The interval is so wide the point estimate carries almost no information. |
|---|---|
| 100 labels | Usable for a single rule. Roughly plus or minus 10 points on a mid-range rate. |
| 400 labels | Roughly halves that width. Each halving costs four times the labels. |
Width shrinks with the square root of the sample, so the fourth hundred labels buys far less than the first hundred. That is the honest reason agreement figures start wide.
Where the uncertainty comes from
Two places, and a serious interval accounts for both. The production sample the rate was measured on, and the much smaller reviewed set the check's own accuracy was measured on.
The second is usually the larger source. A few hundred human labels are standing in for every conversation the system will ever score.
The trap on a rare failure
The interval on recall is driven by how many real failures are in the sample, not by how many conversations are.
At 5 percent prevalence, 400 conversations contain about 20 failures, which is an interval of roughly plus or minus 18 points. Sampling more of the flagged ones and reweighting is the fix, and the weighting has to be stated.
An agreement figure quoted with no sample size and no interval should be read as a claim rather than a measurement. Judge agreement
This page as markdown