# The kappa paradox

> Why two reviewers can agree on 95 percent of conversations and still score a kappa near zero, and what to report instead.

Two reviewers read 1,000 conversations and give the same answer on 940 of them. Kappa comes back at 0.37. Neither number is wrong.

There are two paradoxes here, not one. Feinstein and Cicchetti described both in a pair of 1990 papers in the Journal of Clinical Epidemiology. Almost everyone quotes the first and almost nobody mentions the second.

## One: rarity crushes kappa

Kappa's chance term is the product of the two raters' marginals. When one category holds nearly all the data, that product climbs toward the observed agreement.

```
1,000 conversations, 5% really violate the rule

                 flagged   not flagged
really failed       20          30
really passed       30         920
```

```
Po    = (20 + 920) / 1000                  = 0.940
Pe    = (50*50 + 950*950) / 1000^2         = 0.905

kappa = (0.940 - 0.905) / (1 - 0.905)      = 0.368
```

The subtraction leaves very little and divides it by very little. One cell moving by a few conversations swings the result hard.

Kappa is measuring agreement on the hard cases only. When there are few hard cases the estimate is both small and unstable, which is correct behaviour read as a bad score.

## Two: lopsided raters score better

The second paradox is stranger. Two raters whose marginal totals are unbalanced against each other can score a higher kappa than a balanced pair at the same observed agreement.

So a pair of reviewers who disagree about how often to flag at all can be rewarded for it. This one is almost never mentioned and it is why bias and prevalence indices get reported alongside.

## Does AC1 fix it

No. It removes this artifact by changing the chance model, and introduces prevalence dependence in the opposite direction.

On the table above, AC1 reads 0.934 where kappa reads 0.368. Both describe the same 940 agreements. Neither is wrong, and quoting only the flattering one is the failure.

> **Note:** Reaching for AC1 because kappa reads low is not automatically honest. At 5 percent prevalence AC1 applies almost no chance correction, so it lands near raw percent agreement. Publish both with the prevalence and let a reader see the gap.

## What to report instead

The thing the paradox is really exposing is that agreement on the rare class is invisible inside one headline number. Report it directly.

```
positive specific agreement = 2a / (2a + b + c)

from the table above        = 40 / 100   = 0.40
negative specific agreement = 1840 / 1900 = 0.97
```

One number said 0.94. This pair says the reviewers agree almost perfectly about clean conversations and agree 40 percent of the time about violations. That is the truth the single figure hid.

Report the counts, the prevalence, and the positive specific agreement. Every chance-corrected coefficient is derivable from those, and none of them can be argued with.

---

Source: https://evidova.com/glossary/kappa-paradox
