Why sessions do not need coordinating
Sharing one world between two agents would corrupt both episodes, because the grader reads end state and neither agent produced it alone. Cloning is cheap: on the World Host an open is a reflink clone of the version’s snapshot, measured in milliseconds, and a session costs roughly 200 KB of memory while it lives.Pin one version for the whole suite. Without a pin, a push halfway through
splits your results across two worlds and the mean means nothing.
A shell loop, for a handful of sessions
--no-wait returns as soon as the platform accepts the session, so the loop does not serialize
on provisioning. Poll each id with gateway worlds session status <sessionId> until ready is
true.
Python, with the SDK’s own runner
run_sessions opens one session per task in a thread pool, waits for each to be ready,
hands your agent a bound toolkit, grades the end state, and closes the session warm. Each agent
call runs inside session.trace_context(), so its traces group under that session’s id.
A task that raises scores 0 and records why in
result.error. The tasks that passed keep
their results.
Python, when you want the loop yourself
open_session is the primitive under the runner. Use it directly when the sessions are not
one-per-scenario — a load test, a sweep over models, one session per end user.
Straight HTTP, from any language
POST /api/public/world-sessions is the same call the CLI and the Python client make. Fire N
of them and drive each returned session.
The response carries
sessionId, ready, versionId, instruction, the world’s tools
and the surfaces you asked for. Poll GET /api/public/world-sessions/{sessionId} until
ready, and DELETE the same path when you are done.
One command for a whole suite
1
Run every scenario against a checkout
async (session, task) => void, it calls session.call or session.toolkit(), and the
task’s own grader scores the end state.2
Gate a pull request on the result
worlds ci pushes the checked-out world, runs every task and exits 1 below the threshold.
See Run API worlds in CI for the matrix and the workflow file.3
Or dispatch the run on the platform
bench dispatch starts the same hosted run the Run button starts, parallelism included;
read it back with gateway bench run <runId> --wait. environment run dispatches one runner
job per model and waits for the verdicts.Verify
Count them in the app. The world’s Overview lists every live session with its engine, state, task, API URL and age. The count in the panel header is how many are up right now. Check one at a time.gateway worlds session status <sessionId> answers for a single
session, for when one of N behaves differently from the rest.
Prove they are isolated. Write a record in one session and read it back in another; the
second session should not see it.
Close every session. A loop that opens and never closes holds containers
until the idle reaper finds them 30 minutes later, and pinned sessions are
never reaped at all. Use the Python context manager,
trap ... EXIT in a
shell loop, or DELETE in a finally.Where to go next
- Spin worlds up and down for one session end to end.
- Run API worlds in CI for the workflow file and the gate.
- Simulations for your users for a session per end user.
- Worlds client for reading rollouts and per-task performance afterwards.