Skip to content
RehearsalDocs

Compare agents and runs

Put two agents, or two versions of one agent, side by side. Which runs can be compared, and why runs on different world versions cannot.

Time
5 minutes
Needs:
one or two finished runs

A comparison answers "is version 2 better than version 1?" with the same measures for both. It is only fair when both agents faced the same jobs with the same checks.

What can be compared

You compareAllowedWhy
Agents inside one runYes. This is the best comparisonSame jobs, same world version, same time
Two runs on the same world versionYesSame jobs and the same checks
Two runs on different world versionsNo. The server refuses with 409Each build writes its own jobs and checks. A score on one version and a score on another measure different things

The refusal says: "runs use different world versions; their graders are not comparable".

To compare an agent across two releases of your application, run it on each world version and read each result by itself. Do not subtract one score from the other.

Compare

  1. Open Compare.
  2. Under Run, choose a run. To compare the agents inside it, stop here.
  3. To compare two runs, choose the Second run (optional).

The page shows a table of Measures with the change between the two, and a table Per scenario with the result of each job.

The measures

MeasureBetter isField
Verified successHighersuccess_rate
Episodes with a rule violationLowerhard_violation_episodes
Claimed done but was notLowerfalse_completions
Duplicate side effectsLowerduplicate_effects
Unneeded customer interruptionsLowerunnecessary_interruptions
Blocked tool callsLowerrejected_attempts
Model cost (USD)Lowercost_usd

The page also shows Episodes scored and Recovered after a fault for each side.

Read a comparison

  1. Look at holdout first. It is the score on jobs that neither version was tuned on.
  2. Then look at rule violations and duplicates. A version with a higher success rate and more duplicates is not better for an agent in production.
  3. Check the episode counts. A difference of one episode in ten is noise. Run more repeats before you decide.
  4. Read the jobs that changed. In Per scenario, find the jobs that one version passes and the other fails, and open an episode of each: see Read results.

Good practice

  • Prefer one run with both agents over two runs. Put the earlier version in every run as a control.
  • Use 2 or more repeats.
  • Keep the world version fixed while you compare versions of an agent. Build a new world version when your application changes, and start a new series of comparisons there.

Next

Checked against rehearsal-kit 0.1.2 on 11 October 2026.

Was this page helpful?

Edit this page

On this page

Was this page helpful?

Edit this page