The kappa paradox
Why two reviewers can agree on 95 percent of conversations and still score a kappa near zero, and what to report instead.
Two reviewers read 1,000 conversations and give the same answer on 940 of them. Kappa comes back at 0.37. Neither number is wrong.
There are two paradoxes here, not one. Feinstein and Cicchetti described both in a pair of 1990 papers in the Journal of Clinical Epidemiology. Almost everyone quotes the first and almost nobody mentions the second.
One: rarity crushes kappa
Kappa's chance term is the product of the two raters' marginals. When one category holds nearly all the data, that product climbs toward the observed agreement.
1,000 conversations, 5% really violate the rule
flagged not flagged
really failed 20 30
really passed 30 920Po = (20 + 920) / 1000 = 0.940
Pe = (50*50 + 950*950) / 1000^2 = 0.905
kappa = (0.940 - 0.905) / (1 - 0.905) = 0.368The subtraction leaves very little and divides it by very little. One cell moving by a few conversations swings the result hard.
Kappa is measuring agreement on the hard cases only. When there are few hard cases the estimate is both small and unstable, which is correct behaviour read as a bad score.
Two: lopsided raters score better
The second paradox is stranger. Two raters whose marginal totals are unbalanced against each other can score a higher kappa than a balanced pair at the same observed agreement.
So a pair of reviewers who disagree about how often to flag at all can be rewarded for it. This one is almost never mentioned and it is why bias and prevalence indices get reported alongside.
Does AC1 fix it
No. It removes this artifact by changing the chance model, and introduces prevalence dependence in the opposite direction.
On the table above, AC1 reads 0.934 where kappa reads 0.368. Both describe the same 940 agreements. Neither is wrong, and quoting only the flattering one is the failure.
What to report instead
The thing the paradox is really exposing is that agreement on the rare class is invisible inside one headline number. Report it directly.
positive specific agreement = 2a / (2a + b + c)
from the table above = 40 / 100 = 0.40
negative specific agreement = 1840 / 1900 = 0.97One number said 0.94. This pair says the reviewers agree almost perfectly about clean conversations and agree 40 percent of the time about violations. That is the truth the single figure hid.
Report the counts, the prevalence, and the positive specific agreement. Every chance-corrected coefficient is derivable from those, and none of them can be argued with.
This page as markdown