Skip to content

Why an outside judge

Why the thing scoring your agent should not be the thing that built it.

The company that built your agent has an interest in that agent looking good. And the model family that wrote a reply is the worst reader of whether the reply was any good, because the same blind spot produced both.

Outside your stack

We read finished conversations from whatever produced them, and we own no part of the agent. That is what lets one set of rules run across every platform you use, so two platforms can be compared instead of each grading itself.

Outside any single model

First readingOne model reads the conversation against the rules
Second readingEvery failure it claims is read again by three more judges from three different companies, none of them the first model's family
All three disagreeThe claim is overturned
They splitA fourth and larger model decides, rather than a head count
One judge times outNothing is overturned. Overturning needs all three to have voted
A share of passesGoes to the same second reading, so the work is not only spent on failures

A failure nobody could confirm is not recorded as a failure. It is left unmeasured, with the reason attached. That is the more expensive answer and the honest one.

A judge that has started waving everything through is invisible if you only re-read what it flagged. That is why the passes are sampled too.

Not outside the rules, and that is deliberate

We do not hold your agent to a standard you did not agree to. Every rule is one question in plain words, and the team that runs the agent approves each one and sets what it counts for.

So a score answers to a rule you can read, not to an opinion of ours. It is the half of this that is dependent on you, said plainly rather than dressed up as independence.

Which version of the rules produced a score travels with the score. A number from one version and a number from another are not comparable and should not sit in the same trend.

What we notice without being asked

The other half is independent, and it is what we look for rather than what we score. Nobody configures any of this.

Learned normalEach agent gets its own band per hour of the week, from its own history
A quiet thresholdNo band is published under two hundred readings, so a new agent gets no invented normal
One reading is not enoughA signal needs consecutive readings outside the band, so a single spike fires nothing
Sustained directionEach agent's weekly score is tested for a slide over ten whole weeks
Causes nobody namedEach week the reasons people gave in their own words are grouped, and three that mean the same thing become a cause for your team to accept or refuse

That is the honest position: independent in what we notice, and answerable to you in what we score.

And it is all checkable

None of the above is worth anything on our word. How closely the judge matches your own reviewers is measured per score area and published inside the product, for your workspace and not for a demo. Read how agreement is measured

Read next

  • Agreement with people · How we measure whether our scores match what your reviewers would have said.
  • What gets scored · Why the fast rules run on every conversation and the ones that cost money run on a share.
  • What a score is · What a number here means, what it does not mean, and what you must not use it for.
This page as markdown