Golden dataset
The reviewed set of conversations an evaluation is checked against, how big it needs to be, and how it goes stale without anyone noticing.
A set of conversations people have reviewed by hand, kept as the reference an automated check is measured against.
How big
Per rule, not in total. A hundred reviewed conversations for one rule is a usable floor; a thousand spread across twenty rules is fifty each, which is not.
It also has to contain failures. A sample of a rare failure drawn at random will contain almost none of them, and recall cannot be measured on a set with nothing to recall.
How it goes stale
- The agent changes, and the old conversations stop resembling what it now produces.
- The rule is reworded, and old labels silently answer a different question.
- Traffic shifts to a new channel, language or customer segment.
- Easy cases accumulate, because easy cases are quicker to review.
What a label has to be attached to
A durable coordinate, not a row produced by one scoring run. If labels are anchored to the output of a run, the next run orphans them and the reviewed set silently empties.
Anchor to the rule, the conversation, and which revision of the transcript was read. All three survive a rescore.
This page as markdown