# Cohen's kappa

> The standard chance-corrected measure of agreement between two raters, what its value means, and the one situation where it misleads badly.

Kappa measures agreement between two raters after subtracting the agreement you would expect if both were guessing with their observed habits.

## The formula

```
kappa = (Po - Pe) / (1 - Pe)

Po = share of items both raters put in the same category
Pe = sum over categories of (rater A's share) * (rater B's share)
```

Proposed by Cohen in 1960, in Educational and Psychological Measurement, and still the default anyone will ask you for.

## Reading the value

| | |
|---|---|
| 1 | Perfect agreement. |
| 0 | Exactly what chance predicts. No information. |
| Below 0 | Worse than chance. Usually a label mix-up rather than real disagreement. |

It does not reach minus one in practice. The floor depends on the marginals, and on skewed data it sits close to zero.

Published bands labelling ranges as fair, moderate or substantial are conventions, not findings. The authors of the most cited scale called their own divisions arbitrary in the paper that introduced them.

## Where it misleads

On skewed data. If 95 of 100 conversations pass and both raters agree on 94 of them, kappa can land near zero while raw agreement reads 94 percent.

That is not a bug in the arithmetic. It is kappa correctly reporting that on data this lopsided, agreeing is easy, so agreement carries little evidence.

It is still misleading if you read it as the quality of the rater rather than as the informativeness of the data.

Never quote kappa alone. Quote it with raw agreement and prevalence, which is the combination that makes the number interpretable. [The kappa paradox](/glossary/kappa-paradox)

---

Source: https://evidova.com/glossary/cohens-kappa
