Skip to main content
Open as many sessions of a world as you have questions for it, and run them at the same time. Each session copies the version’s snapshot on open, so two sessions of the same world share nothing: a record written in one is invisible in the other, and a crash in one leaves the rest untouched. A suite is a dozen scenarios, and each one needs a world of its own.

Why sessions do not need coordinating

Sharing one world between two agents would corrupt both episodes, because the grader reads end state and neither agent produced it alone. Cloning is cheap: on the World Host an open is a reflink clone of the version’s snapshot, measured in milliseconds, and a session costs roughly 200 KB of memory while it lives.
Pin one version for the whole suite. Without a pin, a push halfway through splits your results across two worlds and the mean means nothing.

A shell loop, for a handful of sessions

--no-wait returns as soon as the platform accepts the session, so the loop does not serialize on provisioning. Poll each id with gateway worlds session status <sessionId> until ready is true.

Python, with the SDK’s own runner

run_sessions opens one session per task in a thread pool, waits for each to be ready, hands your agent a bound toolkit, grades the end state, and closes the session warm. Each agent call runs inside session.trace_context(), so its traces group under that session’s id.
A task that raises scores 0 and records why in result.error. The tasks that passed keep their results.

Python, when you want the loop yourself

open_session is the primitive under the runner. Use it directly when the sessions are not one-per-scenario — a load test, a sweep over models, one session per end user.
The context manager closes the session on the way out, warm when the block succeeded.

Straight HTTP, from any language

POST /api/public/world-sessions is the same call the CLI and the Python client make. Fire N of them and drive each returned session.
The response carries sessionId, ready, versionId, instruction, the world’s tools and the surfaces you asked for. Poll GET /api/public/world-sessions/{sessionId} until ready, and DELETE the same path when you are done.

One command for a whole suite

1

Run every scenario against a checkout

Each task opens one live session. Your module’s default export is async (session, task) => void, it calls session.call or session.toolkit(), and the task’s own grader scores the end state.
2

Gate a pull request on the result

worlds ci pushes the checked-out world, runs every task and exits 1 below the threshold. See Run API worlds in CI for the matrix and the workflow file.
3

Or dispatch the run on the platform

bench dispatch starts the same hosted run the Run button starts, parallelism included; read it back with gateway bench run <runId> --wait. environment run dispatches one runner job per model and waits for the verdicts.

Verify

Count them in the app. The world’s Overview lists every live session with its engine, state, task, API URL and age. The count in the panel header is how many are up right now. Check one at a time. gateway worlds session status <sessionId> answers for a single session, for when one of N behaves differently from the rest. Prove they are isolated. Write a record in one session and read it back in another; the second session should not see it.
Close every session. A loop that opens and never closes holds containers until the idle reaper finds them 30 minutes later, and pinned sessions are never reaped at all. Use the Python context manager, trap ... EXIT in a shell loop, or DELETE in a finally.

Where to go next