Skip to content
RehearsalDocs

Bring your own agent

Test agent code that you run yourself, in any framework. Rehearsal sends each turn over the API, and your agent answers with one action.

Time
30 minutes to connect an agent
Needs:
a published world version, and a key with Build and evaluate access

Use this when your agent is code: LangGraph, the OpenAI Agents SDK, CrewAI, or your own loop. Rehearsal does not run a model for such an agent. It runs the world, the simulated customer and the checks. Your code decides each action.

How it works

  1. You create an agent profile that is marked as your own.
  2. You start a run with it. The episodes start and wait for your agent.
  3. For each waiting episode, your agent reads a turn: who it plays, what is new, and which tools it may call.
  4. Your agent sends one action. Rehearsal carries it out in the world and returns the result.
  5. This repeats until the episode ends. Then the checks run, as for any agent.

Each turn waits 10 minutes

If no action arrives within 10 minutes of a turn, the episode ends with external_timeout. Start your agent's loop as soon as the run starts.

Set up

Create the profile

Open Agents, choose Create a profile, and turn on My own agent sends actions through the API.

Start a run

rehearsal eval <wv_id> --profile <profile_id> --split dev --repeats 1

Keep the run_id.

Run your agent's loop

Use the Python client, or plain HTTP. Both are below.

The loop in Python

import uuid
from rehearsal.sdk import Rehearsal

rh = Rehearsal()
run = rh.evaluations.create(WORLD_VERSION_ID, [AGENT_ID], splits=["dev"], repeats=1)
memory = {}  # (episode id, actor) -> everything that actor has seen

while rh.runs.get(run["run_id"])["status"] not in ("finished", "failed", "cancelled"):
    for episode in rh.runs.episodes(run["run_id"]):
        if episode["status"] != "running":
            continue
        observation = rh.external.observation(episode["id"])
        if not observation.get("awaiting_action"):
            continue
        turn = observation["turn"]["data"]
        history = memory.setdefault((episode["id"], turn["actor"]), [])
        history += turn["new"]

        tool, args = my_agent_decide(history, turn["tools"])  # your agent: exactly one tool call

        result = rh.external.act(episode["id"], turn["actor"], tool, args, idempotency_key=uuid.uuid4().hex)
        history.append({"role": "tool", "name": tool, "content": str(result)})

my_agent_decide is your code. It gets the conversation so far and the tools, and returns one tool name with its arguments.

  • turn["tools"] uses the OpenAI function format: {"type": "function", "function": {"name", "description", "parameters"}}. Most agent frameworks accept it as their list of tools.
  • turn["new"] has only what arrived since the actor's last action. Keep your own memory for each episode and actor, as the example does.
  • A run can have several episodes waiting at once, one for each sandbox.

The repository has a complete example with a LangGraph ReAct agent: examples/agents/langgraph_react.py.

The loop over HTTP

Every request carries Authorization: Bearer <your_key>.

1. Read the turn

GET/v1/episodes/<episode_id>/observationNeeds access: read
Response
{
  "episode_status": "running",
  "awaiting_action": true,
  "turn": {
    "data": {
      "actor": "mira",
      "new": [{"role": "system", "content": "You are mira (maintainer). ..."}],
      "goal_version_observed": 1,
      "tools": [{"type": "function", "function": {"name": "create_comment", "description": "...", "parameters": {}}}]
    }
  }
}

The actor name and the tool are examples. new holds chat messages: the first turn has the role's instructions, and later turns have tool results, customer messages and messages from other actors. When awaiting_action is false, no turn is waiting: ask again in a second or two.

2. Send one action

POST/v1/episodes/<episode_id>/actionsNeeds access: write
{
  "actor": "mira",
  "tool": "create_comment",
  "args": {"index": 12, "body": "Thank you. The fix is in the next release."},
  "idempotency_key": "6f1c2a9e7b0d4c53"
}

The server answers 202 at once. The action is carried out in the world after that.

3. Get the result

GET/v1/episodes/<episode_id>/actions/<idempotency_key>Needs access: read
Response
{"status": "done", "result": {}}

While the action is in progress, the answer is {"status": "pending"}. Ask again in half a second.

What the idempotency key is for

The key names one action. You choose it: 8 to 64 characters, different for each action. A random UUID is a good key.

  • If your request is lost and you send the same action again with the same key, the server knows it and does not carry it out twice.
  • You use the same key to get the result.
  • The same key with a different action is refused with 409.

Errors from an action

StatusMessageMeaning
409episode is not runningThe episode has ended
409actor does not match the current external turnThe turn belongs to another actor
409current external turn already has an actionOne action for each turn. Read the next turn
409idempotency key already used for a different actionUse a new key for a new action
409episode does not use an external agentThe profile is not a bring-your-own profile
422tool is not offered in the current external turnCall only the tools that the turn lists

Tools that every world offers

Besides the application's own tools, a turn can offer these.

ToolUse
reply_to_customer(text)Talk to the customer. Each contact interrupts them, so keep contacts necessary
send_message(to, text)Message another actor, for example to ask an administrator to make a change that your role cannot
retry_operation(operation_id)Safely send an earlier write again after an UNKNOWN result. It returns the original effect if the write did happen
wait()Wait for new messages, for example a teammate's answer
complete(summary)Say that the customer's latest request is done and checked. Only one actor has it

When the customer talks through the application itself, such as a chat channel or an issue comment, their messages arrive as notifications, and you answer with the application's own tools.

What a careful agent does

Your agent is tested on these. They are also good practice in production.

  1. Read the customer's latest message before every write. Requests change during the work.
  2. When a write returns "not carried out: a new message from the customer arrived", read the new message, then decide again.
  3. After an UNKNOWN result, do not repeat the write blindly. Read the state first, or use retry_operation.
  4. Find ids with read tools. Never guess an id.
  5. Call complete only after a read shows that the work is done.
  6. Do not contact the customer more than necessary.

A key for the agent only

A key with "Build and evaluate" access can do everything on this page. To give agent code less, create a key with only the agent scope. It can read the waiting turn of an episode and send its action, and nothing else. Another program, with its own key, must then start the run and tell the agent the episode ids.

curl -X POST https://rehearsal.example.com/v1/api-keys \
  -H "Authorization: Bearer <a_full_access_key>" -H "Content-Type: application/json" \
  -d '{"name": "agent runner", "scopes": ["agent"]}'

Act as the agent from an AI assistant

An AI assistant can play the agent through two MCP tools: waiting_turns returns the episodes that wait, and take_action sends one action. This is a good way to see how a careful agent behaves, or to check that a job is fair. See the MCP tool reference.

Next

Checked against rehearsal-kit 0.1.2 on 11 October 2026.

Was this page helpful?

Edit this page

On this page

Was this page helpful?

Edit this page