Skip to content
RehearsalDocs

Evaluate an agent

Run one or more agents through the jobs of a published world. Choose the splits and repeats, and know the episode count before you start.

Time
2 to 10 minutes for a run
Needs:
a published world version

An evaluation is a run: a set of agents, each attempting a set of jobs a number of times. Each attempt is an episode, and each episode ends with an outcome from the application's own database.

Before you start

  • You need a world version with the status published. If you have none, build a world.
  • You need an agent profile, which you create below.

Create an agent profile

A profile is one version of an agent: its instructions and its model.

  1. Open Agents and choose Create a profile.
  2. Enter a Name and the Instructions.
  3. Leave Model chain (optional) empty to use the server's default model.
  4. Choose Create profile.

About the instructions:

  • They become the agent's system prompt. Empty instructions use the platform's default.
  • Each job gives each actor its own role instructions. Rehearsal adds them after yours. To place them yourself, put {instructions} in your text. You can also use {actor_id} and {role}.
  • A second profile with the same name becomes version 2. Versions are kept, so results stay comparable.

If your agent is code that you run, not instructions, see Bring your own agent.

Choose what to run

ChoiceAdvice
World versionThe highest published version, unless you compare with an older result
AgentsOne or more. Agents in the same run get the same jobs at the same time, which makes a fair comparison. Keep the earlier version in the run as a control
Splitsdev while you explore and debug. holdout only for the final score, after you have stopped changing the agent
Repeats2 or more. Models are not deterministic, and one attempt at a job is noise

When you choose no split and no repeats, a run uses every approved job, twice.

Count the episodes first

episodes = jobs in the chosen splits x agents x repeats

To count the jobs in each split:

rehearsal worlds scenarios <wv_id>

Only jobs with the status approved are run.

The stress layer

Every episode gets two extra events, on top of what its job contains:

  1. The agent's first change is refused with "too many requests". The change is not carried out. The agent must try it again.
  2. After a change goes through, the customer sends the same request again. The agent must check and answer. It must not do the work twice.

An agent that works only on the easy path fails these.

Start the run

  1. Open Runs and choose Run an evaluation.
  2. Choose the World version and the Agents.
  3. Choose the Scenario splits and the Repeats per scenario.
  4. Choose Start evaluation.

The console opens the run page.

The answer has run_id, episodes (the count) and shards (how many sandboxes run the episodes side by side).

Quiet hours on the Free plan

On the Free plan, a run starts only in the quiet window. The run is accepted and waits. The answer has starts_at, the time the run will start, and the console shows it. Paid plans and your own server start at once.

Follow the run

The run page shows the agents, the customer and the verifier at work, and a table of episodes that fills as they finish.

A run usually takes 2 to 10 minutes. If it stays in progress with no new episode, the workers are busy with other work.

Pause or stop a run

A run is several server jobs, one for each sandbox. A command to the run goes to all of them.

The run page has Pause and Stop.

Episodes that a stopped run never started are not counted against your plan.

When it is refused

MessageMeaningDo this
402: "This run needs ... episodes and the workspace has ... left this month"The plan's episodes do not cover the runRun fewer: one split, fewer jobs, or one repeat
429: "All ... sandboxes of the ... plan are in use by another run"The plan runs a fixed number of sandboxes at onceWait for the other run, or stop it
409The world version is not publishedUse a published version
422No approved job matches the chosen splitsCheck the splits with rehearsal worlds scenarios

Next

Checked against rehearsal-kit 0.1.2 on 11 October 2026.

Was this page helpful?

Edit this page

On this page

Was this page helpful?

Edit this page