Recall on fail
The share of genuinely bad conversations a check actually catches. The number that decides whether a quality system is worth running.
Of the conversations that really did break a rule, how many did the check flag? That is recall on the failure class, and it is the number a quality system lives or dies on.
recall on fail = failures caught / failures that really happenedWhy this one and not overall accuracy
If three percent of conversations fail, a check that flags nothing at all is 97 percent accurate. Accuracy is worthless on skewed data and this data is always skewed.
Recall on fail cannot be gamed that way. A check that flags nothing scores zero, which is the correct description of it.
What it costs to raise
Catching more real failures almost always means raising more false positives. The trade is real and the right balance depends on who reads the flags and what happens next.
State both. A recall figure with no precision beside it hides how much noise was bought to achieve it.