Skip to content

Golden dataset

The reviewed set of conversations an evaluation is checked against, how big it needs to be, and how it goes stale without anyone noticing.

A set of conversations people have reviewed by hand, kept as the reference an automated check is measured against.

How big

Per rule, not in total. A hundred reviewed conversations for one rule is a usable floor; a thousand spread across twenty rules is fifty each, which is not.

It also has to contain failures. A sample of a rare failure drawn at random will contain almost none of them, and recall cannot be measured on a set with nothing to recall.

How it goes stale

  • The agent changes, and the old conversations stop resembling what it now produces.
  • The rule is reworded, and old labels silently answer a different question.
  • Traffic shifts to a new channel, language or customer segment.
  • Easy cases accumulate, because easy cases are quicker to review.

What a label has to be attached to

A durable coordinate, not a row produced by one scoring run. If labels are anchored to the output of a run, the next run orphans them and the reviewed set silently empties.

Anchor to the rule, the conversation, and which revision of the transcript was read. All three survive a rescore.

This page as markdown