Every word we use

The plain phrase for each idea in the product, and the formal name behind it.

One idea, one name. Every word below is a word you will meet on a screen, next to the sentence that says what it means.

The plain phrase is the name. Where a formal or statistical name exists, it comes last on the line, for a reader who wants to see the maths. It is never the label you land on.

ConversationOne unbroken stretch of talking with the agent. If the same customer comes back later, that is a second conversation.
CustomerOne person the agent talked to, however many conversations they have had. Fewer customers than conversations, whenever somebody came back.
ReviewedA person looked at this and recorded a verdict on it.
Quality scoreThe judged score, 0 to 100, for how well the agent handled the conversation.
Score areaOne specific thing we grade a conversation on, phrased as a question a supervisor would ask out loud.
RuleOne specific thing we check for in a conversation.
Automatic failThis rule failing caps the score no matter what else went right, so say what it caps it to, not that a guard was breached.
What we checkThe full list of rules we grade conversations against.
ResultWhat one rule found on one conversation.
Not measuredWe have not measured this yet. The reason lives in the sentence next to it, not the label. A blank is never a zero.
Waiting for approvalThis rule runs and records verdicts, but nothing it decides reaches a score until a reviewer approves it.
StalledThis has stopped producing results and needs attention.
No examples scored by a person yetNobody has hand-labeled a conversation for this yet, so the judge scores it but the result is not a verdict.
AbandonmentThe customer left before the conversation finished.
Could have been savedRecoverable with a faster or clearer reply, ours to fix in how we respond.
Needs a product fixA defect to fix upstream. No reply would have saved this one.
JudgeThe AI that reads a conversation and scores it.
Rules versionWhich version of what-we-check produced this score. A score compares only to another score from the same version.
Two judges agreedHow often independent judges reach the same verdict, once agreement you'd get from luck alone is taken out. For engineers: AC1 (Gwet's AC1). An older, related figure, Cohen's kappa, is shown beside it for comparison; kappa sags when almost everything passes even where the judges really do agree.
Matches peopleHow often the judge's verdict matches what a person would have said, measured against conversations a person hand-scored.
Flag accuracyWhen the judge flags a failure, how often it's a real one. Higher is better: a low number means a lot of false alarms. For engineers: Precision.
Catch rateOf the real failures, how many the judge catches. Higher is better: a low number means a lot of missed problems. For engineers: Recall.
Same conversation, scored twice: usually within ±3 pointsHow much a score moves on its own when nothing about the conversation changed. A change smaller than this is not real movement. For engineers: Coefficient of variation (cv) against a gate; also called the noise floor.
Examples a person scoredConversations a person read and gave a verdict on, with who and when. Precision, recall and agreement are all measured by comparing the judge against these.
Scored by the fast checkThis conversation went through the quick, cheaper pass rather than the full review. Hover for why.
Fully reviewedThis conversation went through the full, slower review.
Spread of scoresHow scores are distributed across all conversations in the period, not just the average.
For engineers: the mathsThe statistics behind the numbers on this page, for someone who wants to check the work.
Screened out before scoringThis conversation never reached the judge, and the specific reason is the message, for example too short to grade, or the customer never replied.
AlertA rule that watches a number and tells you when it crosses a line you set.
Alert firedThe moment an alert's rule was crossed.
Time rangeThe period a number on screen was measured over.
Audit reportA dated report of what we measured over a period, how much we covered, what the results were, and what the team did about the alerts.
SearchFind a conversation by what was said in it, or go to any screen.
Scoring qualityHow good the scoring itself is: how often two judges agree, how closely the judge matches a person, and how much a score moves on a rerun.
Emerging patternA way conversations keep failing that no rule in force asks about yet. It is proposed with the conversations it explains, so you can decide whether it should count.
This page is built from the same file the product reads its labels from, so a word cannot change on a screen and stay old here.

Read next

  • Dimensions and severity · The six things a check can be about, and the four levels of how much it matters.
  • Platform matrix · All twelve platforms side by side: channel, audio, timestamps, delivery, retries.
  • API keys · What a key is, the two scopes, and how to revoke one.
This page as markdown