Why an outside judge

Why the thing scoring your agent should not be the thing that built it.

The company that built your agent has an interest in that agent looking good. And the model family that wrote a reply is the worst possible reader of whether the reply was any good, because the same blind spot produced both. So the scoring here sits outside both of them.

Outside your stack

We read finished conversations from whatever produced them, over an address they post to or a schedule we poll, and we own no part of the agent. That is what lets one set of rules run across every platform you use, so two platforms can be compared instead of each grading itself.

Outside any single model

One model reads the conversation against the rules first. Every failure it claims is then read again, separately, by three more judges from three different companies, and none of the three is the first model's own family.

If all three disagree with the claim, the claim is overturned. If they split, a fourth and larger model decides, rather than a head count deciding. Overturning needs all three to have voted, so one judge timing out cannot reverse anything on its own.

A failure nobody could confirm is not recorded as a failure. It is left unmeasured, with the reason attached. That is the more expensive answer and the honest one.

A share of the passes goes to the same second reading, so the work is not only spent on failures. A judge that has started waving everything through is invisible if you only re-read what it flagged.

Outside our own opinion of the rules

Every rule is one question in plain words, and the rules belong to your workspace. You can add one, retire one, or let a new one run and record what it finds without letting it touch a score until a reviewer approves it.

Which version of the rules produced a score travels with the score, because a number from one version and a number from another are not comparable and should not sit in the same trend.

And it is all checkable

None of the above is worth anything on our word. How closely the judge matches your own reviewers is measured for each score area and published inside the product, for your workspace and not for a demo. That is the next page.

Read next

  • Agreement with people · How we measure whether our scores match what your reviewers would have said.
  • What gets scored · Why the fast rules run on every conversation and the ones that cost money run on a share.
  • What a score is · What a number here means, what it does not mean, and what you must not use it for.
This page as markdown