runSessions runs a regression suite. It opens one container per task, runs your agent against each of them in parallel, grades every session, closes them warm, and returns one report with a pass or fail verdict.
Each task needs its own container: two agents sharing one world state would corrupt both episodes.
Run every scenario and gate the mean
tasks given, every bundled task in the world runs. The world is resolved once for the whole run, so every session pins the same version and a push mid-suite cannot split the results across two worlds.
What the handler receives
Youragent function is called once per task with a session that is already ready and a toolkit already bound to it. It runs inside session.withTraceContext, so the spans it records after tracing.init() group under that session’s id.
Grading, closing and aggregation are handled for you. Return whatever you like from the handler; the reward comes from the task’s grader, not from the return value.
Every option
The report
runSessions resolves to a SessionRunReport.
Each
SessionRunResult carries task, sessionId, reward, rewards, and an error string when the task threw instead of grading.
A task that throws scores zero and records why. The other tasks keep their results.
Mix bundled, stored and inline tasks
tasks accepts all three kinds of reference in the same array.
inline:<hash> in the results.
Stream results as they land
onResult fires as each task finishes.
runSessions runs your agent against real containers. To run the platform’s own agent instead, dispatch a run config with world.dispatch, or gate a pull request with validateWorld.Where to go next
- Gate a pull request when the world source itself is what changed.
- World sessions for the methods on the
sessionyour handler receives. - Run many sessions at once for the same idea from the terminal and from Python.