Skip to content

Recall on fail

The share of genuinely bad conversations a check actually catches. The number that decides whether a quality system is worth running.

Of the conversations that really did break a rule, how many did the check flag? That is recall on the failure class, and it is the number a quality system lives or dies on.

recall on fail = failures caught / failures that really happened

Why this one and not overall accuracy

If three percent of conversations fail, a check that flags nothing at all is 97 percent accurate. Accuracy is worthless on skewed data and this data is always skewed.

Recall on fail cannot be gamed that way. A check that flags nothing scores zero, which is the correct description of it.

What it costs to raise

Catching more real failures almost always means raising more false positives. The trade is real and the right balance depends on who reads the flags and what happens next.

State both. A recall figure with no precision beside it hides how much noise was bought to achieve it.

To know recall at all, somebody has to have labelled a sample by hand, including the conversations the check did not flag. A system that only reviews its own flags can never measure what it missed.
This page as markdown