# Krippendorff's alpha

> The agreement coefficient that handles more than two raters, missing ratings and ordered categories, which kappa cannot.

Alpha measures agreement across any number of raters, over data where not everyone rated everything, and across categories that may be ordered rather than merely different.

## What it handles that kappa does not

| | |
|---|---|
| More than two raters | Kappa is defined for exactly two. |
| Missing ratings | Alpha uses whatever pairs exist per item. |
| Ordered categories | Disagreeing by one level can count as less than disagreeing by three. |
| Different data types | Nominal, ordinal, interval and ratio, through the distance function. |

## How it is computed

From a coincidence matrix rather than a contingency table. Every pairable pair of ratings on an item contributes, weighted so items rated more times do not count more.

```
alpha = 1 - (observed disagreement / expected disagreement)

1   = perfect agreement
0   = agreement at chance
< 0 = systematic disagreement
```

## When to reach for it

When a panel of reviewers rated an overlapping but unequal set of conversations, which is what happens whenever review is done by real people with other work.

For a straight comparison of one automated check against one human reviewer on a complete set, AC1 or kappa is the simpler and more readable answer.

> Alpha is often described as kappa for many raters. It is not. On nominal data with two raters and no gaps it converges on Scott's pi, which pools the raters' marginals instead of keeping them separate, and it inherits the same prevalence problem kappa has.

---

Source: https://evidova.com/glossary/krippendorffs-alpha
