Every word we use
The plain phrase for each idea in the product, and the formal name behind it.
One idea, one name. Every word below is a word you will meet on a screen, next to the sentence that says what it means.
The plain phrase is the name. Where a formal or statistical name exists, it comes last on the line, for a reader who wants to see the maths. It is never the label you land on.
| Conversation | One unbroken stretch of talking with the agent. If the same customer comes back later, that is a second conversation. |
|---|---|
| Customer | One person the agent talked to, however many conversations they have had. Fewer customers than conversations, whenever somebody came back. |
| Reviewed | A person looked at this and recorded a verdict on it. |
| Quality score | The judged score, 0 to 100, for how well the agent handled the conversation. |
| Score area | One specific thing we grade a conversation on, phrased as a question a supervisor would ask out loud. |
| Rule | One specific thing we check for in a conversation. |
| Automatic fail | This rule failing caps the score no matter what else went right, so say what it caps it to, not that a guard was breached. |
| What we check | The full list of rules we grade conversations against. |
| Result | What one rule found on one conversation. |
| Not measured | We have not measured this yet. The reason lives in the sentence next to it, not the label. A blank is never a zero. |
| Waiting for approval | This rule runs and records verdicts, but nothing it decides reaches a score until a reviewer approves it. |
| Stalled | This has stopped producing results and needs attention. |
| No examples scored by a person yet | Nobody has hand-labeled a conversation for this yet, so the judge scores it but the result is not a verdict. |
| Abandonment | The customer left before the conversation finished. |
| Could have been saved | Recoverable with a faster or clearer reply, ours to fix in how we respond. |
| Needs a product fix | A defect to fix upstream. No reply would have saved this one. |
| Judge | The AI that reads a conversation and scores it. |
| Rules version | Which version of what-we-check produced this score. A score compares only to another score from the same version. |
| Two judges agreed | How often independent judges reach the same verdict, once agreement you'd get from luck alone is taken out. For engineers: AC1 (Gwet's AC1). An older, related figure, Cohen's kappa, is shown beside it for comparison; kappa sags when almost everything passes even where the judges really do agree. |
| Matches people | How often the judge's verdict matches what a person would have said, measured against conversations a person hand-scored. |
| Flag accuracy | When the judge flags a failure, how often it's a real one. Higher is better: a low number means a lot of false alarms. For engineers: Precision. |
| Catch rate | Of the real failures, how many the judge catches. Higher is better: a low number means a lot of missed problems. For engineers: Recall. |
| Same conversation, scored twice: usually within ±3 points | How much a score moves on its own when nothing about the conversation changed. A change smaller than this is not real movement. For engineers: Coefficient of variation (cv) against a gate; also called the noise floor. |
| Examples a person scored | Conversations a person read and gave a verdict on, with who and when. Precision, recall and agreement are all measured by comparing the judge against these. |
| Scored by the fast check | This conversation went through the quick, cheaper pass rather than the full review. Hover for why. |
| Fully reviewed | This conversation went through the full, slower review. |
| Spread of scores | How scores are distributed across all conversations in the period, not just the average. |
| For engineers: the maths | The statistics behind the numbers on this page, for someone who wants to check the work. |
| Screened out before scoring | This conversation never reached the judge, and the specific reason is the message, for example too short to grade, or the customer never replied. |
| Alert | A rule that watches a number and tells you when it crosses a line you set. |
| Alert fired | The moment an alert's rule was crossed. |
| Time range | The period a number on screen was measured over. |
| Audit report | A dated report of what we measured over a period, how much we covered, what the results were, and what the team did about the alerts. |
| Search | Find a conversation by what was said in it, or go to any screen. |
| Scoring quality | How good the scoring itself is: how often two judges agree, how closely the judge matches a person, and how much a score moves on a rerun. |
| Emerging pattern | A way conversations keep failing that no rule in force asks about yet. It is proposed with the conversations it explains, so you can decide whether it should count. |
This page is built from the same file the product reads its labels from, so a word cannot change on a screen and stay old here.
Read next
- Dimensions and severity · The six things a check can be about, and the four levels of how much it matters.
- Platform matrix · All twelve platforms side by side: channel, audio, timestamps, delivery, retries.
- API keys · What a key is, the two scopes, and how to revoke one.