Skip to main content
runSessions runs a regression suite. It opens one container per task, runs your agent against each of them in parallel, grades every session, closes them warm, and returns one report with a pass or fail verdict. Each task needs its own container: two agents sharing one world state would corrupt both episodes.

Run every scenario and gate the mean

With no tasks given, every bundled task in the world runs. The world is resolved once for the whole run, so every session pins the same version and a push mid-suite cannot split the results across two worlds.

What the handler receives

Your agent function is called once per task with a session that is already ready and a toolkit already bound to it. It runs inside session.withTraceContext, so the spans it records after tracing.init() group under that session’s id. Grading, closing and aggregation are handled for you. Return whatever you like from the handler; the reward comes from the task’s grader, not from the return value.

Every option

The report

runSessions resolves to a SessionRunReport. Each SessionRunResult carries task, sessionId, reward, rewards, and an error string when the task threw instead of grading.
A task that throws scores zero and records why. The other tasks keep their results.

Mix bundled, stored and inline tasks

tasks accepts all three kinds of reference in the same array.
An unnamed inline task is labelled inline:<hash> in the results.

Stream results as they land

onResult fires as each task finishes.
runSessions runs your agent against real containers. To run the platform’s own agent instead, dispatch a run config with world.dispatch, or gate a pull request with validateWorld.

Where to go next