Skip to content
RehearsalDocs

MCP tool reference

The 18 tools of the Rehearsal MCP server. For each one, its arguments, what it returns, and whether it spends money.

An AI assistant drives Rehearsal through these tools. The server runs on your machine, started by the assistant with rehearsal mcp, and it uses the sign-in that rehearsal login saved. Every call has the access of that key and nothing more. To install the tools, see AI assistants.

Tools that spend money

4 of the 18 tools start work that spends model budget from the workspace's monthly allowance. The assistant must say what it will start and what it will cost, and wait for a clear yes. Every other tool is free.

ToolTypical costTypical time
build_world$3-8 (stops itself at its budget)1-3 hours
recheck_worldabout $0.03 per job5-20 minutes
run_evaluationabout $0.02 per episode (jobs x agents x repeats)2-10 minutes
improve_agentunder $0.10about 1 minute

The costs are the product's own estimates, in the skill that ships with this version. A plan limit is a hard stop whatever the assistant decides: see What costs money.

What the server tells the assistant

The server gives every assistant these instructions when it connects:

Rehearsal tests AI agents on a sandboxed copy of a real application, with fictional data, a simulated customer who changes their mind, and injected faults; results are checked against the application's own database. Typical flow: list_applications or add_application -> build_world (1-3 hours; check with job_status) -> list_worlds -> list_agents or create_agent -> run_evaluation -> run_results -> episode_details for failures -> improve_agent -> run_evaluation again on the holdout split -> compare_runs. Builds and evaluations spend model money from the workspace's monthly allowance: say what you are about to start and ask the user before build_world, run_evaluation, improve_agent and recheck_world. To act as the agent under test yourself, create_agent(bring_your_own=True), start run_evaluation with it, then loop: waiting_turns -> read the new messages -> take_action with one of the offered tools.

Errors

A tool that fails returns the server's message with its HTTP status. 402 means the plan does not cover the request: do not retry it. 429 means something is still in progress: wait. When no one is signed in, every tool answers that the user must run rehearsal login. See Errors.

Workspace and applications

Check the connection and the plan, and see or add the applications that worlds are built from.

workspace_status

The workspace's plan and what it has left this month (applications, world builds, episodes, sandboxes), model spend against the monthly allowance, and what exists in the workspace.

workspace_status()

Returns: plan (its name, its limits, what is used, the renewal date, the quiet hours, a link to the Billing page), usage (model spend this month against the allowance), the number of applications, published world versions and agents, and a console link.

The same from elsewhere: CLI rehearsal whoami; Python rh.get("/v1/usage").

list_applications

Applications in the workspace (the real apps that practice worlds are built from).

list_applications()

Returns: A list. Each item has id, name, source (repository and commit) and notes.

The same from elsewhere: CLI rehearsal apps list; Python rh.applications.list().

add_application

Add an application from a public git repository. images maps service name to a pinned image (for example {"app": "org/app:1.2@sha256:..."}). policies are rules agents must follow, in plain words.

add_application(name, repository_url, commit="", images=None, policies="")
ArgumentTypeDefaultMeaning
namestringrequiredA name for the application, unique in the workspace.
repository_urlstringrequiredAddress of a public git repository.
commitstring""The full commit SHA to pin.
imagesobjectnoneEach service name with its pinned image, for example {"app": "org/app:1.2@sha256:..."}.
policiesstring""Rules the agent must follow, in plain words. Jobs turn them into hard rules.

Returns: id and name of the new application.

The same from elsewhere: CLI rehearsal apps add; Python rh.applications.create(name, source, images, policies).

Worlds

Build a practice world, follow the build, and see the worlds and their jobs.

build_world

Spends money: ask first

Start compiling an application into a practice world. Takes 1-3 hours and spends model money (about $3-8). Ask the user first. Follow progress with job_status.

build_world(application_id, scenarios=8, budget_usd=None)
ArgumentTypeDefaultMeaning
application_idstringrequiredFrom list_applications.
scenariosinteger8How many jobs to write. Fewer is faster and costs less.
budget_usdnumbernoneModel spend at which the build stops itself.

Returns: job_id, world_id, world_version_id, version, and watch: a console link to the live build.

Cost: $3-8 (stops itself at its budget). Time: 1-3 hours.

The same from elsewhere: CLI rehearsal build; Python rh.worlds.build(application_id).

job_status

Status of a build, evaluation shard or improvement job: state, finished steps, spend, and the error if it failed.

job_status(job_id)
ArgumentTypeDefaultMeaning
job_idstringrequiredFrom build_world, run_evaluation, improve_agent or recheck_world.

Returns: id, kind, status (queued, running, paused, succeeded, failed or cancelled), attempts, spent_usd, completed_steps, result, and error (shortened) when the job failed.

The same from elsewhere: CLI rehearsal jobs show; Python rh.jobs.get(job_id).

control_job

pause, resume, stop, or retry (restart a failed or stopped job from its last saved step).

control_job(job_id, command)
ArgumentTypeDefaultMeaning
job_idstringrequiredThe job to control.
commandstringrequiredpause, resume, stop or retry.

Returns: The job after the command.

The same from elsewhere: CLI rehearsal jobs pause; Python rh.jobs.command(job_id, command).

list_worlds

Practice worlds and their versions. Only published versions can be used in evaluations.

list_worlds()

Returns: A list of worlds. Each has world, application_id and versions; each version has world_version_id, version and status.

The same from elsewhere: CLI rehearsal worlds list; Python rh.worlds.list().

list_scenarios

The jobs (scenarios) in a world version, with their split: train and dev are for improving, holdout is for the final score.

list_scenarios(world_version_id)
ArgumentTypeDefaultMeaning
world_version_idstringrequiredFrom list_worlds.

Returns: A list of jobs. Each has id, slug, title, split and status. Only approved jobs are evaluated.

The same from elsewhere: CLI rehearsal worlds scenarios; Python rh.worlds.scenarios(world_version_id).

recheck_world

Spends money: ask first

Replay every job of a published world and try each with a live agent; unfair jobs are withdrawn. Costs about $0.03 per job. Ask first.

recheck_world(world_version_id)
ArgumentTypeDefaultMeaning
world_version_idstringrequiredA published world version.

Returns: A job. Follow it with job_status.

Cost: about $0.03 per job. Time: 5-20 minutes.

The same from elsewhere: CLI rehearsal worlds validate; Python rh.worlds.validate(world_version_id).

Agents and evaluations

Create agent versions, run them through jobs, read the results, and improve them.

list_agents

Agent profiles (versions of an agent: instructions, tool guidance, model).

list_agents()

Returns: A list. Each item has id, name, version, driver (llm: Rehearsal runs the model; external: you bring the agent) and created_by.

The same from elsewhere: CLI rehearsal profiles list; Python rh.profiles.list().

create_agent

Create an agent profile. instructions is its system prompt (empty uses the platform default). With bring_your_own=True, Rehearsal does not run a model for it: your own agent acts through waiting_turns and take_action.

create_agent(name, instructions="", model=None, bring_your_own=False)
ArgumentTypeDefaultMeaning
namestringrequiredName of the agent. A second profile with the same name becomes version 2.
instructionsstring""The agent's system prompt. Empty uses the platform default.
modelstringnoneA model chain on the server. Empty uses the server's default.
bring_your_ownbooleanfalsetrue when your own agent acts through waiting_turns and take_action.

Returns: id, name and version of the new profile.

The same from elsewhere: CLI rehearsal profiles create; Python rh.profiles.create(name, system_prompt).

run_evaluation

Spends money: ask first

Run agents through the jobs of a published world version. By default every job (train, dev, holdout) is run twice, with the default stress layer on. Narrow it with splits (any of train, dev, holdout) and repeats. Spends model money (about $0.02 per episode). Ask the user first.

run_evaluation(world_version_id, agent_ids, splits=None, repeats=None, budget_usd=None)
ArgumentTypeDefaultMeaning
world_version_idstringrequiredA published world version, from list_worlds.
agent_idslist of stringsrequiredOne or more agents, from list_agents. Several agents in one run get the same jobs.
splitslist of stringsnoneAny of train, dev, holdout. Without it, every split.
repeatsintegernoneAttempts at each job. Without it, 2.
budget_usdnumbernoneModel spend at which the run stops itself.

Returns: run_id, job_id, episodes, shards, and watch: a console link to the live run. On the free plan, starts_at when the run waits for quiet hours.

Cost: about $0.02 per episode (jobs x agents x repeats). Time: 2-10 minutes.

The same from elsewhere: CLI rehearsal eval; Python rh.evaluations.create(world_version_id, profile_ids).

run_results

Status and scores of an evaluation run, with one line per episode.

run_results(run_id)
ArgumentTypeDefaultMeaning
run_idstringrequiredFrom run_evaluation.

Returns: status, summary (scores for each agent), one line for each episode (id, split, scenario, status, verified, ending), and a console link.

The same from elsewhere: CLI rehearsal runs show; Python rh.runs.get(run_id).

episode_details

What happened in one episode: the request and its revisions, every action with its result, and which checks failed.

episode_details(episode_id)
ArgumentTypeDefaultMeaning
episode_idstringrequiredFrom run_results.

Returns: scenario, goals (the request and its revisions), status, ending, verified, environment_error, failed_checks (each with check, observed, wanted), duplicate_effects, false_completion, actions, the last 30 messages, and a console link.

The same from elsewhere: CLI rehearsal runs episode; Python rh.runs.episode(episode_id).

improve_agent

Spends money: ask first

Propose the next version of an agent from its failed train and dev episodes in a run. Ask first. Then evaluate both versions on the holdout split and compare_runs before keeping the new one.

improve_agent(agent_id, run_id)
ArgumentTypeDefaultMeaning
agent_idstringrequiredThe agent to improve.
run_idstringrequiredA finished run with train or dev episodes of that agent.

Returns: A job. When it succeeds, its result names the new profile version and gives the reason for each change.

Cost: under $0.10. Time: about 1 minute.

The same from elsewhere: CLI rehearsal improve; Python rh.post("/v1/improvements", {...}).

compare_runs

Compare agents within one run, or two runs on the same world version.

compare_runs(run_a, run_b=None)
ArgumentTypeDefaultMeaning
run_astringrequiredA run.
run_bstringnoneA second run on the same world version. Without it, the agents inside run_a are compared.

Returns: The scores of each agent, side by side.

The same from elsewhere: CLI rehearsal compare; Python rh.evaluations.compare(run_a, run_b).

Bring your own agent

Act as the agent under test: read each turn, then send one action.

waiting_turns

For a run of a bring-your-own agent: episodes waiting for your next action. Each has the actor you play, the new messages since your last action, and the tools you may call. Waits up to wait_s for a turn.

waiting_turns(run_id, wait_s=30)
ArgumentTypeDefaultMeaning
run_idstringrequiredA run of a bring-your-own agent.
wait_sinteger30How long to wait for a turn, in seconds. At most 120.

Returns: A list of episodes that wait for your action. Each has episode_id, actor (who you play), new_messages (since your last action) and tools (each with name, description and parameters as JSON Schema). An empty list means no turn is waiting.

The same from elsewhere: Python rh.external.observation(episode_id).

take_action

Act in an episode as the actor named by waiting_turns, with one of the tools it offered. Returns the tool's result.

take_action(episode_id, actor, tool, arguments=None)
ArgumentTypeDefaultMeaning
episode_idstringrequiredFrom waiting_turns.
actorstringrequiredThe actor named by waiting_turns.
toolstringrequiredOne of the tools that turn offered.
argumentsobjectnoneThe tool's arguments.

Returns: The result of the tool call.

The same from elsewhere: Python rh.external.act(episode_id, actor, tool, args, idempotency_key).

Generated from the code of rehearsal-kit 0.1.2.

Was this page helpful?

On this page

Was this page helpful?