Guides
One task on each page, in the order you usually do them, from adding an application to comparing two versions of an agent.
Each guide does one task and shows it in the console, from a terminal, in Python and through an AI assistant. The guides follow the order of the work.
Applications and worlds
Add an applicationWhat your application must have, and how to pin its source and images.Build a practice worldThe nine phases, how long they take, and what to do when a build fails.
Agents
Evaluate an agentChoose jobs, splits and repeats, and know the episode count first.Read resultsWhat each outcome, ending and number means, and how to read a failed check.Improve an agentTurn failures into a better version, and check it on jobs it never saw.Compare agents and runsWhich runs can be compared, and why.
Your own agent
Automate
The usual path
- Add an application, once for each application.
- Build a practice world, once for each release of the application that you want to test against.
- Evaluate an agent on the
devjobs, and read the results. - Improve the agent, by hand or with the improvement step.
- Evaluate the old and the new version on the
holdoutjobs, and compare them.
Checked against rehearsal-kit 0.1.2 on 11 October 2026.
Was this page helpful?