# Testing Retell agents before release

> Rehearse a change against calls that already happened, so a prompt edit is measured on real traffic before it meets a caller.

Most testing of a voice agent is simulated: write a scenario, run a synthetic caller, see what happens. It catches the obvious break and misses the thing your actual callers do.

Your own history is the better test set, because it is the only one that contains the questions people really ask.

## Rehearse against calls you already have

1. Import your Retell history.
   Backfilled calls are scored by the same rules as everything arriving live.
2. Write the rule the change is supposed to satisfy.
   A change nobody can state as a rule is a change nobody can prove worked.
3. Try the rule on past calls before it goes live.
   You see which calls it would have failed, and read them, before a single score depends on it.

## Then prove the change landed

After the release, compare the same rule over the same window and the same kind of traffic. A fix that moves the number is a measurement. A fix that nobody measured is a feeling.

Every score is frozen under the version of the rules it was scored with, so a later edit to a rule cannot quietly rewrite last month.

## Where this sits beside Retell's own testing

Retell has simulation and AI QA, and they cover the before and the sampled after. This covers the whole of what shipped, with the rule you can edit and the transcript under every result.

Nothing about your agent changes. We read the call Retell has already finished.

---

Source: https://evidova.com/platforms/retell/testing
