# F1 score

> The harmonic mean of precision and recall, and the specific situations where averaging the two into one number is the wrong thing to do.

F1 combines precision and recall into one figure, using the harmonic mean so that a low value in either one drags the result down.

```
F1 = 2 * (precision * recall) / (precision + recall)
```

A plain average would let perfect precision hide terrible recall. The harmonic mean does not: it sits closer to the smaller of the two.

## When it is the wrong summary

Whenever the two errors cost different amounts, which for conversation quality is nearly always. F1 treats a false alarm and a missed failure as equally bad.

A missed compliance breach and a wasted minute of a reviewer's time are not equally bad, and a metric that says they are will push you to the wrong operating point.

It also ignores true negatives entirely, so it says nothing about how quiet a check is on the conversations that were fine.

## What to do instead

Quote precision and recall separately and let the reader apply their own cost. One number is convenient for a leaderboard and lossy for a decision.

If one number is unavoidable, the Matthews correlation coefficient is the better choice. It uses all four cells of the table, so unlike F1 it cannot ignore how often a clean conversation was wrongly flagged.

> F1 is often called balanced or robust to class imbalance. It is neither. It discards true negatives entirely and moves with prevalence, so two F1 scores from datasets with different failure rates are not comparable. Chicco and Jurman set this out in 2020.

---

Source: https://evidova.com/glossary/f1-score
