> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Worlds client

> Open live world sessions, store tasks, stream rows in, and read a world's scenarios, per-task performance and rollouts — the gatewaysdk clients for worlds and the runs made against them.

Read and drive [worlds](/worlds) from Python. Four modules divide the work: `world_sessions` opens a live copy an agent can call, `world_tasks` keeps the questions you ask it, `world_data` puts rows in, and `environments` reads back the **scenarios** (the tasks a world is asked), the **rollouts** (single graded attempts), and per-task **performance**.

All four talk to the same Surface Area host over the public REST API, and pair with the [Worlds hub](/sdk/benchmark-hub) client that pushes and pins versions.

To create or execute a populated Slack world, follow [Build a Slack world](/worlds/slack).

<Info>
  The product says **world**; the SDK module, CLI namespace and REST paths say
  `environment`. The words on screen moved to the six in the
  [Glossary](/glossary) while identifiers stayed put, so existing integrations
  keep working. Read `environment` in code as `world`.
</Info>

<Info>
  The client reads `GATEWAY_HOST`, `GATEWAY_PUBLIC_KEY`, and
  `GATEWAY_SECRET_KEY` from the environment (same as the benchmark push client).
  Reads accept the public key; writes (`create_task`, `trigger_agent`) require
  the **secret key**.
</Info>

## Open a live session and drive it

`open_session` boots (or reuses) a copy of the world pinned to one version, opened on one task. The call returns as soon as the platform accepts the session, so wait for `ready()` before the first tool call.

```python theme={null}
from gatewaysdk.world_sessions import open_session

with open_session("acme-crm", "refund-double-charge", surfaces=["api"]) as session:
    session.ready()
    print(session.instruction)          # what the agent is told to do
    print(session.api, session.api_headers)

    tools = session.toolkit()           # the task's tools, as callables
    session.call("list_accounts", {"limit": 10})

    print(session.grade())              # {"reward": ..., "rewards": {...}, "raw": ...}
```

| Member | What it does |
| - | - |
| `ready(timeout=...)` | Block until the process inside the container has asked for work. |
| `instruction`, `task_info`, `version_id` | What the session was opened on, and the exact version it pins. |
| `call(tool, args)`, `toolkit()` | Make one tool call, or get the task's tools as Python callables. |
| `api`, `api_headers`, `ui`, `browser()` | The surfaces asked for at open time. |
| `native_computer(browser, provider)` | The browser as Claude's or OpenAI's own computer-use tool. See [Native computer use](/worlds/computer-use). |
| `seed(rows, mode=...)` | Put rows into the live world through its contract. |
| `state()`, `export()`, `calls()` | The world's current rows; the call log. |
| `grade(timeout=...)` | Run the task's grader against the current state. |
| `reset()` | Back to the opening state — base data plus the task's seed. |
| `close(keep_warm=False)` | End the session, or hand the container back warm. |

`export()` holds what the platform brokered. The world's own request log is the `calls` a grader reads: every call the world answered, its route requests (including the ones your agent made through the browser or straight to `api.url`) and its tools called by name, each with its `via` and with credential fields read as `[redacted]`. Read it from the live session, before you close it:

```python theme={null}
from gatewaysdk.world_request_log import WorldRequestLog

log = WorldRequestLog.read(session)          # None when the world keeps no request log
for request in (log.requests if log else []):
    print(request.get("seq"), request.get("via"), request.get("method"), request.get("path"), request.get("status"))
lines = log.as_calls() if log else []        # the same requests as call-log entries with kind "request"
```

`open_session(slug, task, *, version_id=None, reuse_warm=True, surfaces=None, browser_tier=None, host=None, public_key=None, secret_key=None)` takes `task` as a bundled scenario's name, a stored task reference `{"id": ..., "version": ...}`, or an inline spec. `surfaces` asks for more than `tools`: `"api"` for the world's HTTP API, `"ui"` for its dashboard, `"browser"` for a hosted browser driving that dashboard. Asking for a surface the world does not have is refused, not downgraded. `browser_tier` picks the capacity behind a World Host browser: `"standard"` (spot, may be reclaimed) or `"premium"` (on-demand); left as `None`, your organization's default applies, and any other value raises `ValueError` before a request is sent.

A `standard` browser can be reclaimed mid-session, never the world. `browser()` reconnects on its own within 90 seconds and reopens the page it was on. A read (`text()`, `screenshot()`, `title()`, `html()`) or a `goto()` is then retried once; a `click()`, `type()`, `press()`, `back()` or `scroll()` is not replayed, since it may already have reached the world, and raises `BrowserReconnected` (with `outcome` and `restored_url`) so the agent looks at the page before acting again. If you drive your own Playwright, `session.repair_browser()` asks the platform for a working browser behind the same `wsUrl` and answers `{"outcome": "healthy" | "replaced" | "starting", "retryAfterMs": ...}`.

## Keep tasks outside the bundle

`gatewaysdk.world_tasks` stores a task — instruction, seed and grader — as a project resource, so many sessions open on it by id and every grade attributes to the task and its version.

```python theme={null}
from gatewaysdk import world_tasks

task = world_tasks.create_task(
    {"instruction": "Refund the duplicate charge on order #4471."},
    name="Refund a double charge",
    world="acme-crm",
)
world_tasks.validate_task(task.id)            # entities checked before any session opens
```

The module also exposes `update_task` (a new spec becomes the next version), `list_tasks`, `get_task`, and `delete_task`.

### Tasks that need several worlds

A task whose spec has a `worlds` table — `{alias: {slug, ref?, seed?, grader?, tools?}}` — is one instruction over several hosted worlds. `open_task` brings every one of them up in one platform call and hands back the sessions keyed by alias; if any world fails to open, the ones that did are closed and the error names it.

```python theme={null}
from gatewaysdk import world_tasks

task = world_tasks.create_task(
    {
        "instruction": "Find the stolen credentials and lock the account.",
        "worlds": {
            "spycloud": {
                "slug": "spycloud-world",
                "seed": {"rows": {"breaches": [{"id": "b1", "email": "a@x.io"}]}},
                "grader": {"kind": "assertions", "checks": [{"entity": "breaches", "where": {"id": "b1"}}]},
            },
            "torchlight": {"slug": "torchlight-world", "ref": "v2"},
        },
    },
    name="orion-infostealer",
)

with world_tasks.open_task(task.id, surfaces=["api"]) as live:
    live.ready()
    my_agent(live.instruction, live.sessions)      # {"spycloud": WorldSession, "torchlight": WorldSession}
    grade = live.grade()                            # TaskGrade(reward, rewards, worlds, ungraded)
```

`open_task(ref, *, version=None, surfaces=None, reuse_warm=True, browser_tier=None, host=None, public_key=None, secret_key=None)` takes a task id, `{"id": ..., "version"?: ...}` or a spec with `worlds`; `browser_tier` applies to every world. The `TaskSessions` it returns carries `task_id`, `task_version`, `name`, `instruction`, `sessions` and `aliases`, plus `session(alias)`, `ready(timeout=None)`, `status()`, `manifest()` (what `gateway worlds task up` prints), `grade(timeout=None, report=None)`, `export()` and `close(keep_warm=False)`, each fanned out over every world. `grade()` means the worlds and the task's cross-world verifiers that answered a number and lists the others under `ungraded` (a verifier as `x.<name>`), never counting a missing reward as 0; each verifier's own grade is under `TaskGrade.verifiers` by name. `grade()` names the task's other sessions on every world's grade, which is how the platform gathers a verifier's evidence — `WorldSession.grade(worlds={alias: session_id})` does the same for a session you grade by hand. See [Verifiers that span worlds](/worlds/sessions#verifiers-that-span-worlds). `attach_task(task_id)` rebuilds the sessions a stored task has up (refusing when an alias is up twice), `attach_manifest(manifest)` rebuilds them from a saved manifest, and `list_task_sessions(task_id)` lists them. Open a multi-world task with `open_task`.

### Run an agent against a multi-world task end to end

`run_task` is the engine behind `gateway worlds task run`: open every world, call your agent once, grade
every world with its report, file one rollout and close — in one call:

```python theme={null}
from gatewaysdk import TaskRun, run_task

def agent(task: TaskRun) -> dict:
    url, token = task.api("spycloud")           # (url, token) for a world opened with surfaces=["api"]
    tools = task.toolkit()                       # every world's tools merged, named "<alias>.<tool>"
    task.record({"visited": list(task.sessions)})
    ...
    return {"report": "locked the account after finding the credential in spycloud"}

summary = run_task(
    "wt_1",                # or {"id": ..., "version": ...}, or a spec with `worlds`
    agent,
    model="claude-sonnet-4-5",
    out_dir="runs",
)
print(summary["reward"], summary["evaluationId"])
```

`agent` is called as `agent(task_run)` with a `TaskRun` — `.run_id`, `.instruction`, `.model`, `.sessions` (alias ->
`WorldSession`), `.manifest`, `.api(alias)`, `.toolkit()`, `.record(dict)` — and its return value is the report: a
`str`, or a `dict` with a `"report"` key (any other keys are kept on the returned `run.json` dict as-is). A coroutine
it returns is awaited via `asyncio.run`. The report is passed to every world's `grade(report=...)` so a rubric judge
reads it as the agent's own account.

`run_task(ref, agent, *, model, surfaces=None, out_dir="runs", trace=True, file=True, keep_up=False, reuse_warm=True, browser_tier=None, host=None, public_key=None, secret_key=None)`
writes `<out_dir>/<run id>/{manifest,grades,run}.json` (and `report.md` when the agent gave one), where `<run id>` is
`wtr-<12 hex>`. Every step is a flag with a default, nothing implicit: tracing runs with the credentials the run's
requests use (`host`/`public_key`/`secret_key`, then `GATEWAY_*`, then the saved `gateway auth login`; a
`GATEWAY_OTLP_ENDPOINT` on another host only with the environment's own keys) unless `trace=False` (a tracing
failure never fails the run — `run.json["traced"]` says whether it actually ran, and the process prints one warning
saying why, never a key), filing runs unless `file=False`, and every session closes
unless `keep_up=True` — even when the agent raises, so grading and filing still happen with the error recorded in
`run.json["agentError"]` and the filed evaluation marked FAILED. Filing reuses `benchmark_hub.evals` exactly like
`gatewaysdk.worlds.World.run` does for a single in-process world, anchored to the task's first world's version (the
platform pins one evaluation to one version; every world's own version and reward travel in the sample's
`metadata`/`info`). A world whose tools the platform could not read (a `tool-contract-unread` notice) raises
`gatewaysdk.world_notices.WorldToolsUnread` (its `code` is `tool-contract-unread`) before the agent runs: every
session closes, even with `keep_up=True`, and nothing is filed. `run_sessions` records such a task as failed without
calling its agent.

## Stream rows into a world

`gatewaysdk.world_data` uploads rows as a validated batch: gzipped chunks the platform checks against the world's contract, applies onto its snapshot, and turns into a data-only version.

```python theme={null}
from gatewaysdk.world_data import Flags, WorldDataClient, import_rows, wait

client = WorldDataClient.from_env()
manifest = import_rows(
    client,
    "acme-crm",
    ["rows.jsonl"],
    flags=Flags(entity="accounts", mode="append", publish="now"),
)
status = wait(client, "acme-crm", manifest.batch_id)
print(status["state"], client.counts("acme-crm"))
```

`Flags` carries the options the CLI spells out: `mode` (`append` upserts by primary key, `replace` swaps named entities), `merge`, `atomic`, `publish` (`now` or `later`), `dryRun`, `entity` for bare-row files, and the chunk sizes. `import_rows` returns a manifest and `wait` polls the batch to a terminal state.

A refused row names the field that broke the contract and leaves the world unchanged. The rows file contract and the batch lifecycle are covered on [Put data in a world](/worlds/data).

## Read a world's results

```python theme={null}
from gatewaysdk.environments import GatewayEnvironmentClient

client = GatewayEnvironmentClient.from_env()  # host + keys from env vars

for env in client.list_environments().environments:
    print(env.slug, env.latestVersion, env.taskSetDatasetId)

perf = client.get_performance(container_id)
for task in perf.tasks:  # worst-first; nulls kept as their own category
    print(task.task, task.avgReward, task.nullScoreCount)
```

Every method returns a validated pydantic model (see
`gatewaysdk.environments.types`). Null scores are always surfaced as their own
category: `avgReward` / `avgScore` are computed over non-null values only, and
`nullScoreCount` tells you how many samples had no score. They are never coerced
to 0.

## Client methods

| Method | Description |
| - | - |
| `list_environments(slug=None, limit=50)` | List the project's environments (id, slug, name, `latestVersion`, `versionCount`, `taskSetDatasetId`). |
| `get_environment(container_id)` | One environment with its version history (`author`, `changeReason`, `envKind`) and rollout count. |
| `get_task_set(container_id)` | The linked `kind='tasks'` dataset and its tasks. 404 if the world has no scenarios. |
| `get_performance(container_id)` | Per-task performance across rollouts, with a per-version breakdown. |
| `list_rollouts(container_id)` | The world's rollouts, newest first, with metrics and sample counts. |
| `get_rollout_samples(container_id, evaluation_id, limit=200)` | The scored samples of one rollout (reward/score/correct + named sub-rewards + info). |
| `create_task(dataset_id, name, prompt, expected_output=None, metadata=None)` | Create a task in a task set (a `kind='tasks'` dataset item). **Write: needs the secret key.** |
| `get_agent()` | The project's Environment Agent config + recent runs. |
| `trigger_agent()` | Enqueue a manual Environment Agent run. **Write: needs a saved config and the secret key.** |

### Creating a task

A task lives in a **task set**, which is a dataset of kind `tasks`. Get its id
from the environment (`get_environment(...).taskSetDatasetId`) or from
`get_task_set(...).datasetId`:

```python theme={null}
env = client.get_environment(container_id)
task = client.create_task(
    env.taskSetDatasetId,
    name="Refund a double charge",
    prompt="The customer was charged twice for order #4471. Issue a refund.",
    expected_output="A single refund of the duplicate charge is issued.",
    metadata={"difficulty": "medium"},
)
print(task.id)
```

The task's author is attributed to the API key used for the write.

### Environment Agent

The Environment Agent is a scheduled analysis agent that reads per-task
performance and can auto-heal weak tasks. Read its status and trigger a run:

```python theme={null}
status = client.get_agent()
if status.config and status.config.enabled:
    run = client.trigger_agent()
    print(run.runId, run.status)
```

`trigger_agent()` requires a saved config: configure the agent in the app
first.

## CLI

The same reads are exposed under `gateway environment` in the [gateway CLI](/cli). Sessions, stored tasks and data have their own groups — see [gateway worlds](/cli/worlds).

```bash theme={null}
gateway environment list
gateway environment get <container-id>
gateway environment tasks <container-id>
gateway environment performance <container-id>
gateway environment rollouts <container-id>
gateway environment samples <container-id> <evaluation-id> --limit 100
gateway environment create-task <dataset-id> --name "..." --prompt "..."
gateway environment agent-status
gateway environment agent-trigger
```

Each command reads `GATEWAY_HOST` / `GATEWAY_PUBLIC_KEY` / `GATEWAY_SECRET_KEY`
(or `--host`) and prints the JSON response.

## REST API

The client is a thin wrapper over these public REST endpoints (project-scoped,
HTTP Basic with your public + secret key):

| Method & path | Description |
| - | - |
| `GET /api/public/environments` | List worlds (`?slug=`, `?limit=`). |
| `GET /api/public/environments/{containerId}` | World detail plus its versions. |
| `GET /api/public/environments/{containerId}/task-set` | Linked task set + tasks (404 if unlinked). |
| `GET /api/public/environments/{containerId}/performance` | Per-task performance. |
| `GET /api/public/environments/{containerId}/rollouts` | Rollouts (evaluations) list. |
| `GET /api/public/environments/{containerId}/rollouts/{evaluationId}/samples` | Rollout samples (`?limit=`). |
| `POST /api/public/task-sets/{datasetId}/tasks` | Create a task (`{ name, prompt, expectedOutput?, metadata? }`). |
| `GET /api/public/environment-agent` | Environment Agent config + recent runs. |
| `POST /api/public/environment-agent/trigger` | Enqueue a manual Environment Agent run. |

Reads accept a public key; the two writes require the secret key.

The session, stored-task and data modules sit on their own routes — `/api/public/world-sessions`, `/api/public/world-tasks` and `/api/public/worlds/{slug}/data/*`. See the [REST API](/rest-api/resources#worlds).

## Where to go next

* [gateway worlds](/cli/worlds) — the same surface from a terminal.
* [Worlds hub](/sdk/benchmark-hub) — pushing and pinning a world's versions from Python.
* [Put data in a world](/worlds/data) — the rows file contract in full.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.