Measurement vocabulary
The statistics behind an AI agent quality number, in plain English, including the ones that mislead.
The words behind a quality number, defined plainly. Most of these have running code under them in this product rather than a definition we looked up, and where the field disagrees about a term, the page says so instead of picking a side quietly.
What this is not: a glossary of voice AI. Barge-in, endpointing and word error rate are explained well elsewhere and do not need a worse page from us.
Agreement
Gwet's AC1A chance-corrected agreement coefficient that stays stable when almost every item falls in one category, which is where Cohen's kappa collapses.Cohen's kappaThe standard chance-corrected measure of agreement between two raters, what its value means, and the one situation where it misleads badly.The kappa paradoxWhy two reviewers can agree on 95 percent of conversations and still score a kappa near zero, and what to report instead.Percent agreementThe simplest agreement measure, why it always looks good on skewed data, and why no serious reliability claim rests on it alone.Krippendorff's alphaThe agreement coefficient that handles more than two raters, missing ratings and ordered categories, which kappa cannot.Judge agreementHow far an automated judge matches the humans it stands in for, which statistic to quote, and what a number without a method is worth.
Classifier
Precision and recallThe two numbers behind any automated quality check, what each one is blind to, and why quoting only one of them hides the failure that matters.Recall on failThe share of genuinely bad conversations a check actually catches. The number that decides whether a quality system is worth running.Confusion matrixThe four outcomes of any pass or fail check, named for conversation quality, and which cell each headline metric is computed from.F1 scoreThe harmonic mean of precision and recall, and the specific situations where averaging the two into one number is the wrong thing to do.PrevalenceHow often the thing you are looking for actually happens, and why a rare failure wrecks precision no matter how good the check is.Confidence intervalThe range a measured rate could really sit in, how sample size moves it, and how many labels an agreement figure needs before it means anything.
Data
Ground truthThe human judgement an automated score is measured against, and why calling it truth hides the most important question about it.Golden datasetThe reviewed set of conversations an evaluation is checked against, how big it needs to be, and how it goes stale without anyone noticing.Label noiseDisagreement and error inside your own human labels, which puts a hard ceiling on the agreement any automated judge can be measured at.LLM as a judgeUsing a language model to score another model's output, what it is good at, and the three failure modes that make an unchecked judge worthless.
Outcome rates
Containment rateThe share of conversations an agent handled without a human, what it does not tell you, and why it rises when an agent gets worse.Deflection rateConversations kept away from a human queue. The most quoted and least defined number in support automation, and what to pin down before you quote it.Resolution rateThe share of conversations ending resolved, the two ways vendors count one, and why the figure on your invoice is a different set.Task completion rateWhether the customer got the thing they came for, how it differs from containment, and why it is the harder number to fake.Escalation rateHow often an agent hands a conversation to a person, why a low one is not automatically good, and what to read beside it.
For the words this product uses on screen, rather than the words the field uses, read every word we use. Every word we use
This page as markdown