Skip to main content
Read and drive worlds from Python. Four modules divide the work: world_sessions opens a live copy an agent can call, world_tasks keeps the questions you ask it, world_data puts rows in, and environments reads back the scenarios (the tasks a world is asked), the rollouts (single graded attempts), and per-task performance. All four talk to the same Surface Area host over the public REST API, and pair with the Worlds hub client that pushes and pins versions. To create or execute a populated Slack world, follow Build a Slack world.
The product says world; the SDK module, CLI namespace and REST paths say environment. The words on screen moved to the six in the Glossary while identifiers stayed put, so existing integrations keep working. Read environment in code as world.
The client reads GATEWAY_HOST, GATEWAY_PUBLIC_KEY, and GATEWAY_SECRET_KEY from the environment (same as the benchmark push client). Reads accept the public key; writes (create_task, trigger_agent) require the secret key.

Open a live session and drive it

open_session boots (or reuses) a copy of the world pinned to one version, opened on one task. The call returns as soon as the platform accepts the session, so wait for ready() before the first tool call.
export() holds what the platform brokered. The world’s own request log is the calls a grader reads: every call the world answered, its route requests (including the ones your agent made through the browser or straight to api.url) and its tools called by name, each with its via and with credential fields read as [redacted]. Read it from the live session, before you close it:
open_session(slug, task, *, version_id=None, reuse_warm=True, surfaces=None, browser_tier=None, host=None, public_key=None, secret_key=None) takes task as a bundled scenario’s name, a stored task reference {"id": ..., "version": ...}, or an inline spec. surfaces asks for more than tools: "api" for the world’s HTTP API, "ui" for its dashboard, "browser" for a hosted browser driving that dashboard. Asking for a surface the world does not have is refused, not downgraded. browser_tier picks the capacity behind a World Host browser: "standard" (spot, may be reclaimed) or "premium" (on-demand); left as None, your organization’s default applies, and any other value raises ValueError before a request is sent. A standard browser can be reclaimed mid-session, never the world. browser() reconnects on its own within 90 seconds and reopens the page it was on. A read (text(), screenshot(), title(), html()) or a goto() is then retried once; a click(), type(), press(), back() or scroll() is not replayed, since it may already have reached the world, and raises BrowserReconnected (with outcome and restored_url) so the agent looks at the page before acting again. If you drive your own Playwright, session.repair_browser() asks the platform for a working browser behind the same wsUrl and answers {"outcome": "healthy" | "replaced" | "starting", "retryAfterMs": ...}.

Keep tasks outside the bundle

gatewaysdk.world_tasks stores a task — instruction, seed and grader — as a project resource, so many sessions open on it by id and every grade attributes to the task and its version.
The module also exposes update_task (a new spec becomes the next version), list_tasks, get_task, and delete_task.

Tasks that need several worlds

A task whose spec has a worlds table — {alias: {slug, ref?, seed?, grader?, tools?}} — is one instruction over several hosted worlds. open_task brings every one of them up in one platform call and hands back the sessions keyed by alias; if any world fails to open, the ones that did are closed and the error names it.
open_task(ref, *, version=None, surfaces=None, reuse_warm=True, browser_tier=None, host=None, public_key=None, secret_key=None) takes a task id, {"id": ..., "version"?: ...} or a spec with worlds; browser_tier applies to every world. The TaskSessions it returns carries task_id, task_version, name, instruction, sessions and aliases, plus session(alias), ready(timeout=None), status(), manifest() (what gateway worlds task up prints), grade(timeout=None, report=None), export() and close(keep_warm=False), each fanned out over every world. grade() means the worlds and the task’s cross-world verifiers that answered a number and lists the others under ungraded (a verifier as x.<name>), never counting a missing reward as 0; each verifier’s own grade is under TaskGrade.verifiers by name. grade() names the task’s other sessions on every world’s grade, which is how the platform gathers a verifier’s evidence — WorldSession.grade(worlds={alias: session_id}) does the same for a session you grade by hand. See Verifiers that span worlds. attach_task(task_id) rebuilds the sessions a stored task has up (refusing when an alias is up twice), attach_manifest(manifest) rebuilds them from a saved manifest, and list_task_sessions(task_id) lists them. Open a multi-world task with open_task.

Run an agent against a multi-world task end to end

run_task is the engine behind gateway worlds task run: open every world, call your agent once, grade every world with its report, file one rollout and close — in one call:
agent is called as agent(task_run) with a TaskRun — .run_id, .instruction, .model, .sessions (alias -> WorldSession), .manifest, .api(alias), .toolkit(), .record(dict) — and its return value is the report: a str, or a dict with a "report" key (any other keys are kept on the returned run.json dict as-is). A coroutine it returns is awaited via asyncio.run. The report is passed to every world’s grade(report=...) so a rubric judge reads it as the agent’s own account. run_task(ref, agent, *, model, surfaces=None, out_dir="runs", trace=True, file=True, keep_up=False, reuse_warm=True, browser_tier=None, host=None, public_key=None, secret_key=None) writes <out_dir>/<run id>/{manifest,grades,run}.json (and report.md when the agent gave one), where <run id> is wtr-<12 hex>. Every step is a flag with a default, nothing implicit: tracing runs with the credentials the run’s requests use (host/public_key/secret_key, then GATEWAY_*, then the saved gateway auth login; a GATEWAY_OTLP_ENDPOINT on another host only with the environment’s own keys) unless trace=False (a tracing failure never fails the run — run.json["traced"] says whether it actually ran, and the process prints one warning saying why, never a key), filing runs unless file=False, and every session closes unless keep_up=True — even when the agent raises, so grading and filing still happen with the error recorded in run.json["agentError"] and the filed evaluation marked FAILED. Filing reuses benchmark_hub.evals exactly like gatewaysdk.worlds.World.run does for a single in-process world, anchored to the task’s first world’s version (the platform pins one evaluation to one version; every world’s own version and reward travel in the sample’s metadata/info). A world whose tools the platform could not read (a tool-contract-unread notice) raises gatewaysdk.world_notices.WorldToolsUnread (its code is tool-contract-unread) before the agent runs: every session closes, even with keep_up=True, and nothing is filed. run_sessions records such a task as failed without calling its agent.

Stream rows into a world

gatewaysdk.world_data uploads rows as a validated batch: gzipped chunks the platform checks against the world’s contract, applies onto its snapshot, and turns into a data-only version.
Flags carries the options the CLI spells out: mode (append upserts by primary key, replace swaps named entities), merge, atomic, publish (now or later), dryRun, entity for bare-row files, and the chunk sizes. import_rows returns a manifest and wait polls the batch to a terminal state. A refused row names the field that broke the contract and leaves the world unchanged. The rows file contract and the batch lifecycle are covered on Put data in a world.

Read a world’s results

Every method returns a validated pydantic model (see gatewaysdk.environments.types). Null scores are always surfaced as their own category: avgReward / avgScore are computed over non-null values only, and nullScoreCount tells you how many samples had no score. They are never coerced to 0.

Client methods

Creating a task

A task lives in a task set, which is a dataset of kind tasks. Get its id from the environment (get_environment(...).taskSetDatasetId) or from get_task_set(...).datasetId:
The task’s author is attributed to the API key used for the write.

Environment Agent

The Environment Agent is a scheduled analysis agent that reads per-task performance and can auto-heal weak tasks. Read its status and trigger a run:
trigger_agent() requires a saved config: configure the agent in the app first.

CLI

The same reads are exposed under gateway environment in the gateway CLI. Sessions, stored tasks and data have their own groups — see gateway worlds.
Each command reads GATEWAY_HOST / GATEWAY_PUBLIC_KEY / GATEWAY_SECRET_KEY (or --host) and prints the JSON response.

REST API

The client is a thin wrapper over these public REST endpoints (project-scoped, HTTP Basic with your public + secret key): Reads accept a public key; the two writes require the secret key. The session, stored-task and data modules sit on their own routes — /api/public/world-sessions, /api/public/world-tasks and /api/public/worlds/{slug}/data/*. See the REST API.

Where to go next