Gwet's AC1
A chance-corrected agreement coefficient that stays stable when almost every item falls in one category, which is where Cohen's kappa collapses.
AC1 measures how far two raters agree beyond what you would expect from luck. It answers the same question as Cohen's kappa and answers it differently when one category dominates.
The formula
AC1 = (Pa - Pe) / (1 - Pe)
Pa = share of items both raters put in the same category
Pe = sum over categories of pi*(1 - pi), divided by (q - 1)
pi = mean of the two raters' marginal share of category i
q = number of categoriesThe numerator is the same as kappa's. The whole difference is the chance term, and the difference matters more than it looks.
What changes against kappa
Kappa estimates chance agreement by multiplying the two raters' marginals. When 95 percent of conversations pass, that product is near 0.9, so almost all the observed agreement is explained away as luck.
AC1 instead weights each category by pi*(1 - pi), which is largest when a category is near half the data and smallest when it is near all or none of it.
In plain terms: AC1 assumes agreeing on a near-universal category is easy and discounts it gently, rather than assuming most of it was a coincidence.
The catch nobody mentions
As the thing you are looking for gets rarer, the AC1 chance term shrinks toward zero, and AC1 converges on plain percent agreement.
At 5 percent prevalence it applies almost no correction at all. Exactly where a chance correction is most needed, AC1 supplies least of it.
Published criticism puts it plainly: AC1 is prevalence-dependent in the opposite direction to kappa, and its zero point is not statistical independence. Swapping to it because kappa reads low trades a pessimistic bias for an optimistic one.
Where it comes from
| Proposed | Gwet, 2008, in the British Journal of Mathematical and Statistical Psychology. |
|---|---|
| Criticised | Vach and Gerke, 2023, in MethodsX, under the title that AC1 is not a substitute for Cohen's kappa. |
The least disputed part of the original paper is its variance estimator, which does not assume the two raters are independent. That is a real contribution whatever you make of the coefficient.
Why this product reports it
Conversation quality is skewed by nature. Most conversations do not break most rules, so a pass-heavy distribution is the normal case rather than an edge case.
Read the kappa paradox next for why the two disagree, and percent agreement for why neither can be replaced by the raw number. The kappa paradox
This page as markdown