> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Run many sessions at once

> Every session is its own clone of a pinned world version, so N sessions run side by side without leaking state. Open them from a shell loop, from Python, or from a CI job.

Open as many sessions of a world as you have questions for it, and run them at the same time.

Each session copies the version's snapshot on open, so two sessions of the same world share
nothing: a record written in one is invisible in the other, and a crash in one leaves the rest
untouched. A suite is a dozen [scenarios](/glossary#scenario), and each one needs a world of
its own.

## Why sessions do not need coordinating

Sharing one world between two agents would corrupt both episodes, because the grader reads end
state and neither agent produced it alone. Cloning is cheap: on the World Host an open is a
reflink clone of the version's snapshot, measured in milliseconds, and a session costs roughly
200 KB of memory while it lives.

<Info>
  Pin one version for the whole suite. Without a pin, a push halfway through
  splits your results across two worlds and the mean means nothing.
</Info>

## A shell loop, for a handful of sessions

```bash theme={null}
export GATEWAY_HOST=https://withgateway.ai
export GATEWAY_PUBLIC_KEY=pk-lf-...
export GATEWAY_SECRET_KEY=sk-lf-...

VERSION=$(gateway bench resolve vendor-world | awk '/^versionId:/ {print $2}')

for task in refund-settled refund-disputed refund-partial; do
  gateway worlds session open vendor-world \
    --task "$task" --surface api --version-id "$VERSION" --no-wait &
done
wait
```

`--no-wait` returns as soon as the platform accepts the session, so the loop does not serialize
on provisioning. Poll each id with `gateway worlds session status <sessionId>` until `ready` is
true.

## Python, with the SDK's own runner

`run_sessions` opens one session per task in a thread pool, waits for each to be ready,
hands your agent a bound toolkit, grades the end state, and closes the session warm. Each agent
call runs inside `session.trace_context()`, so its traces group under that session's id.

```python theme={null}
import os

from gatewaysdk.world_sessions import run_sessions

# GATEWAY_HOST, GATEWAY_PUBLIC_KEY and GATEWAY_SECRET_KEY come from the environment.
assert os.environ.get("GATEWAY_SECRET_KEY"), "set GATEWAY_SECRET_KEY"


def my_agent(session, toolkit, instruction, task):
    """Drive one world. `toolkit.impls` are callables bound to this session."""
    accounts = toolkit.impls["list_accounts"]()
    if accounts:
        toolkit.impls["close_account"](id=accounts[0]["id"])


report = run_sessions(
    "vendor-world",
    lambda session, toolkit, instruction, task: my_agent(
        session, toolkit, instruction, task
    ),
    concurrency=4,
    fail_under=0.7,
)

print(f"mean reward {report.mean:.2f}, passed {report.passed}")
for result in report.results:
    print(result.task, result.reward, result.error or "")
```

| Argument | Default | What it does |
| - | - | - |
| `tasks` | every task in the world | The scenarios to run; bundled names or inline specs |
| `concurrency` | `3` | How many sessions run at once |
| `version_id` | resolved once | Pins every session to the same version |
| `fail_under` | `None` | Sets `report.passed` from the mean reward |
| `keep_warm` | `True` | Returns each container to the warm pool on close |
| `on_result` | `None` | Called with each `SessionRunResult` as it finishes |

A task that raises scores 0 and records why in `result.error`. The tasks that passed keep
their results.

## Python, when you want the loop yourself

`open_session` is the primitive under the runner. Use it directly when the sessions are not
one-per-scenario — a load test, a sweep over models, one session per end user.

```python theme={null}
import os
from concurrent.futures import ThreadPoolExecutor

from gatewaysdk.world_sessions import open_session

MODELS = ["openai/gpt-5-mini", "anthropic/claude-sonnet-4-5"]


def attempt(model: str) -> dict:
    with open_session("vendor-world", "refund-settled", surfaces=["api"]) as session:
        session.ready()
        base_url = session.api                 # this session's own HTTP API
        headers = session.api_headers          # the bearer header it needs
        run_my_agent(model, base_url, headers, session.instruction)
        return {"model": model, **session.grade()}


with ThreadPoolExecutor(max_workers=len(MODELS)) as pool:
    for graded in pool.map(attempt, MODELS):
        print(graded["model"], graded.get("reward"))
```

The context manager closes the session on the way out, warm when the block succeeded.

## Straight HTTP, from any language

`POST /api/public/world-sessions` is the same call the CLI and the Python client make. Fire N
of them and drive each returned session.

```bash theme={null}
curl -sS -X POST "$GATEWAY_HOST/api/public/world-sessions" \
  -u "$GATEWAY_PUBLIC_KEY:$GATEWAY_SECRET_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
        "slug": "vendor-world",
        "task": "refund-settled",
        "surfaces": ["api"],
        "reuseWarm": true
      }'
```

| Body field | Type | Notes |
| - | - | - |
| `slug` | string | The world; resolved to its head READY version |
| `versionId` | string | Pin an exact version instead of `slug` |
| `task` | string or object | A bundled scenario's name, `{"id": ...}` for a stored task, or an inline spec |
| `surfaces` | array | Besides `tools`: `api`, `ui`, `browser` |
| `reuseWarm` | boolean | Reuse an idle container for the same version and task |
| `browserTier` | string | `standard` (spot, may be reclaimed) or `premium` (on-demand) for a World Host browser; default: the organization's setting |

The response carries `sessionId`, `ready`, `versionId`, `instruction`, the world's `tools`
and the `surfaces` you asked for. Poll `GET /api/public/world-sessions/{sessionId}` until
`ready`, and `DELETE` the same path when you are done.

## One command for a whole suite

<Steps>
  <Step title="Run every scenario against a checkout">
    ```bash theme={null}
    gateway worlds run ./vendor-world --agent ./agent.mjs --all
    gateway worlds run vendor-world@main --agent ./agent.ts --task refund-settled
    ```

    Each task opens one live session. Your module's default export is
    `async (session, task) => void`, it calls `session.call` or `session.toolkit()`, and the
    task's own grader scores the end state.
  </Step>

  <Step title="Gate a pull request on the result">
    ```bash theme={null}
    gateway worlds ci ./vendor-world --agent ./agent.mjs --min-score 0.7
    ```

    `worlds ci` pushes the checked-out world, runs every task and exits 1 below the threshold.
    See [Run API worlds in CI](/worlds/ci) for the matrix and the workflow file.
  </Step>

  <Step title="Or dispatch the run on the platform">
    ```bash theme={null}
    gateway bench dispatch <runConfigId>          # ids: world page -> Runs -> Configs
    gateway environment run vendor-world -m openai/gpt-5-mini
    ```

    `bench dispatch` starts the same hosted run the **Run** button starts, parallelism included;
    read it back with `gateway bench run <runId> --wait`. `environment run` dispatches one runner
    job per model and waits for the verdicts.
  </Step>
</Steps>

## Verify

**Count them in the app.** The world's Overview lists every live session with its engine,
state, task, API URL and age. The count in the panel header is how many are up right now.

**Check one at a time.** `gateway worlds session status <sessionId>` answers for a single
session, for when one of N behaves differently from the rest.

**Prove they are isolated.** Write a record in one session and read it back in another; the
second session should not see it.

```bash theme={null}
gateway worlds session state "$SESSION_A"    # JSON: every entity's rows, as the world holds them
gateway worlds session state "$SESSION_B"
```

<Info>
  **Close every session.** A loop that opens and never closes holds containers
  until the idle reaper finds them 30 minutes later, and pinned sessions are
  never reaped at all. Use the Python context manager, `trap ... EXIT` in a
  shell loop, or `DELETE` in a `finally`.
</Info>

## Where to go next

* [Spin worlds up and down](/worlds/sessions) for one session end to end.
* [Run API worlds in CI](/worlds/ci) for the workflow file and the gate.
* [Simulations for your users](/worlds/for-your-users) for a session per end user.
* [Worlds client](/sdk/environments) for reading rollouts and per-task performance afterwards.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.