F1 score
The harmonic mean of precision and recall, and the specific situations where averaging the two into one number is the wrong thing to do.
F1 combines precision and recall into one figure, using the harmonic mean so that a low value in either one drags the result down.
F1 = 2 * (precision * recall) / (precision + recall)A plain average would let perfect precision hide terrible recall. The harmonic mean does not: it sits closer to the smaller of the two.
When it is the wrong summary
Whenever the two errors cost different amounts, which for conversation quality is nearly always. F1 treats a false alarm and a missed failure as equally bad.
A missed compliance breach and a wasted minute of a reviewer's time are not equally bad, and a metric that says they are will push you to the wrong operating point.
It also ignores true negatives entirely, so it says nothing about how quiet a check is on the conversations that were fine.
What to do instead
Quote precision and recall separately and let the reader apply their own cost. One number is convenient for a leaderboard and lossy for a decision.
If one number is unavoidable, the Matthews correlation coefficient is the better choice. It uses all four cells of the table, so unlike F1 it cannot ignore how often a clean conversation was wrongly flagged.