Precision and recall
The two numbers behind any automated quality check, what each one is blind to, and why quoting only one of them hides the failure that matters.
Any check that flags a conversation is a classifier, whether it is a model or a regular expression. These are the two numbers that describe it.
precision = correct flags / all flags raised
recall = correct flags / all real failuresWhat each one cannot see
| Precision | Blind to failures you never flagged. It only scores the flags you raised. |
|---|---|
| Recall | Blind to false positives. Flag everything and recall is perfect. |
This is why one alone is never a claim. A check that flags nothing has undefined precision and zero recall. One that flags everything has perfect recall and useless precision.
They also behave differently when the failure rate changes. Recall is unaffected by it. Precision moves with it, even when the check has not changed at all.
So a precision figure measured on a balanced test set does not survive contact with production, and a vendor quoting one without its prevalence is quoting a number you cannot use.
Which one to optimise
It depends on what a mistake costs. If a flagged conversation goes to a person to read, a false positive costs their time and precision matters.
If a missed failure means a customer was told something untrue and nobody found out, recall matters more and it is not close.
For conversation quality the asymmetry is usually severe, which is why recall on the failure class is the number worth arguing about. Recall on fail
This page as markdown