Skip to content

Krippendorff's alpha

The agreement coefficient that handles more than two raters, missing ratings and ordered categories, which kappa cannot.

Alpha measures agreement across any number of raters, over data where not everyone rated everything, and across categories that may be ordered rather than merely different.

What it handles that kappa does not

More than two ratersKappa is defined for exactly two.
Missing ratingsAlpha uses whatever pairs exist per item.
Ordered categoriesDisagreeing by one level can count as less than disagreeing by three.
Different data typesNominal, ordinal, interval and ratio, through the distance function.

How it is computed

From a coincidence matrix rather than a contingency table. Every pairable pair of ratings on an item contributes, weighted so items rated more times do not count more.

alpha = 1 - (observed disagreement / expected disagreement)

1   = perfect agreement
0   = agreement at chance
< 0 = systematic disagreement

When to reach for it

When a panel of reviewers rated an overlapping but unequal set of conversations, which is what happens whenever review is done by real people with other work.

For a straight comparison of one automated check against one human reviewer on a complete set, AC1 or kappa is the simpler and more readable answer.

Alpha is often described as kappa for many raters. It is not. On nominal data with two raters and no gaps it converges on Scott's pi, which pools the raters' marginals instead of keeping them separate, and it inherits the same prevalence problem kappa has.
This page as markdown