# Golden dataset

> The reviewed set of conversations an evaluation is checked against, how big it needs to be, and how it goes stale without anyone noticing.

A set of conversations people have reviewed by hand, kept as the reference an automated check is measured against.

## How big

Per rule, not in total. A hundred reviewed conversations for one rule is a usable floor; a thousand spread across twenty rules is fifty each, which is not.

It also has to contain failures. A sample of a rare failure drawn at random will contain almost none of them, and recall cannot be measured on a set with nothing to recall.

## How it goes stale

- The agent changes, and the old conversations stop resembling what it now produces.
- The rule is reworded, and old labels silently answer a different question.
- Traffic shifts to a new channel, language or customer segment.
- Easy cases accumulate, because easy cases are quicker to review.

## What a label has to be attached to

A durable coordinate, not a row produced by one scoring run. If labels are anchored to the output of a run, the next run orphans them and the reviewed set silently empties.

Anchor to the rule, the conversation, and which revision of the transcript was read. All three survive a rescore.

---

Source: https://evidova.com/glossary/golden-dataset
