gateway worlds covers a world end to end: its contract, its rows, its versions, the live sessions an agent calls, and the tests it ships. The groups below follow that order.
Every command reads credentials the way the CLI overview describes, and --host overrides the host on any of them.
Run
gateway worlds <group> --help. The help text is generated from the
command tree, so it matches the build you have installed.Author a world
A world starts from a shipped connector template, a connector your organization published, or a blank contract you write yourself.worlds create when the contract already exists. It is the same request as POST /api/public/worlds and the MCP tool create_world, and it answers with the world plus its first READY version.
url, the world’s page in the app (https://<host>/project/<projectId>/world?env=<containerId>). --json prints the whole answer, whose links carry api (the world’s JSON descriptor, https://<host>/world/<slug>), world (the same descriptor at /api/public/worlds/<slug>), tasks and sessions.
Its flags: --from (template slug, workspace:<slug>, or a directory), --overlay (files laid over the source before compiling), --rows (repeatable seed file) with --entity for bare-row files, --name, --description, --message/-m, and --json. --batch <file> reads JSONL lines of {slug, from, overlay?, data?, name?} and prints one result line each, exiting 2 when any line failed.
The world is stored as its schema tree
The platform stores thegateway-world/1 tree as you wrote it, with one session serving both its tools and its HTTP routes.
Shape the contract
The contract inschema/world.json declares entities, fields, keys and relationships. connector.toml declares the vendor connection and the routes the world answers. Two commands edit them without hand-editing the files.
custom scaffold ships one example entity, records, and schema init lists gateway worlds schema entity remove <dir> records among its next steps. Declare your own entities first, then remove the example:
connector.toml authenticates with a bearer token, the kind a hosted session hands the agent. gateway worlds serve answers 401 to a request without one and accepts the bearer --api-key pins (conform when you pass none).
The entity must exist before a route can name it. Declare write routes too: a lookup-only mock tests nothing about what an agent does when it creates, updates and polls.
Check what you wrote
conform measures fidelity. Its failure list is what to fix to make the replica behave like the vendor.
reconcile works through that list. Vendor documents drift from live accounts: snake_case where the docs say camelCase, undocumented keys, required fields the account omits, timestamps without a zone. Reconciling against captures/ fixes those without hand-editing. It changes top-level fields only, prints every change it makes, and writes nothing until --write. It runs on a local python3 3.12 or newer.
Capture data from the real vendor
Two connector commands turn live vendor responses into a world’s starting data.--remote runs the capture on the platform instead — from the pod whose address the vendor allow-listed, using the credential held as a project secret — and pulls the redacted pages back into captures/ here. --environment names the world for that path and --timeout bounds the wait.
--dry-run reports every refusal at once, grouped by capture, entity and pointer, and exits 2 when there is one. --report needs --dry-run, and writes the same records data import --report does plus the capture each came from.
A capture makes real calls against a real account. Scope the credential, name
the one operation you want, and cap the pages before running it.
Put rows into a world
worlds data is the way in for rows that do not come from a capture. A rows file is {"<entity>": [row, ...]} JSON, {entity, row} JSONL, or bare rows with --entity naming what they are.
<world> argument decides where the rows land. A directory writes into the tree. A slug streams the rows up as gzipped chunks into a data batch that the platform validates, applies onto the world’s snapshot, and turns into a version.
import flags worth knowing: --mode append|replace (append upserts by primary key), --merge upsert|fill inside the batch, --redact off|apply|refuse (default off), --publish now|later, --atomic (one refusal refuses the batch, the default) against --per-chunk, --wait/--no-wait, --resume to pick up an interrupted upload, --chunk-rows/--chunk-bytes/--parallel for large files, --report for the refusals as JSONL, and --dry-run.
--redact off|apply|refuse decides what happens to plaintext on the way in, and data import, data check, data ingest and session seed all take it. Redaction is optional: off is the default and rows land as given. Keep real data out covers apply and refuse.
A refused row names the field that broke the contract and leaves the world
unchanged.
--report refused.jsonl works on import and check alike.Publish a connector template
worlds connector publish <path> saves the template as a workspace connector — created on the first publish, a new commit after — so other projects can start worlds from the same contract.
A world directory is mapped back into the template layout on publish;
data/, captures/, db/ and generated files stay behind. The published connector equals the directory, so a path you removed is removed there too.
Publishing the world — the versioned bundle a session runs — is gateway bench push. See bench & versions.
Open live sessions
A session is one live copy of the world, pinned to a version, opened on a task. The agent calls it over the surfaces it asked for.open first prints session <id> opening — gateway worlds session status <id> on stderr, so a wait cut short keeps the id. It then prints {sessionId, ready, surfaces, api.url, ui.entry, browser.wsUrl, task, idleExpiresAt, pinned}. Ask for surfaces beyond tools with --surface/-s (api, ui, browser, repeatable or comma-separated); the platform answers 400 rather than quietly serving fewer than you asked for. Pin an exact version with --version-id, and pass --no-wait to return as soon as the platform accepts the session. Otherwise open waits for ready up to --timeout seconds (600 by default): a session that fails to open exits 1 with the reason and its code, and one still opening when the time runs out is printed with ready: false and exits 1 while it keeps opening.
A World Host session’s
api block also carries a token. Send it as
Authorization: Bearer <token> in a header on every request to api.url —
never in the URL.task block reports which was used.
A session closes on its own after thirty idle minutes. Pinning keeps its URL up until you close it. A closed session can come back with
reopen: its checkpoint is the platform’s last export of it while it ran.
Keep tasks as project resources
A stored task is an instruction, a seed and a grader kept by the project rather than baked into a bundle, so many sessions can open on it by id.validate catches a grader asserting on an entity the world does not define, which would otherwise pass against an empty result. For a task whose spec names its worlds, validate checks each alias against the ref it names and answers checked: "worlds" with a per-alias result.
Tasks that need several worlds
A spec with aworlds table — {alias: {slug, ref?, seed?, grader?, tools?}} — is one instruction over several hosted worlds. These commands take a task id or the manifest task up --out wrote; see Spin worlds up and down.
A check that needs more than one world goes in the spec’s verifiers list: each has a name, a scope of two or more aliases, and an assertions or python grader that reads rows as <alias>.<entity> and call logs as <alias>.calls. task create, task update --spec and session open --task-file in the npm CLI check the file before sending: an alias outside worlds or listed twice, a one-alias scope, a rubric kind, a repeated name, a repeated check id, a key that does not belong to the verifier’s kind, a check entity without an alias from the scope, or a first scope alias whose grader is a rubric exits 2 with nothing sent. From the Python CLI, the platform catches the same mistakes and the command exits 1. Without worlds, a verifier runs on the session of its first scope alias, so the spec’s name must be that alias. A verifier’s reward shows up in task grade as x.<name>; see Verifiers that span worlds.
Run an agent against a multi-world task
worlds task run does up, one call to your agent, grade, filing and down in a single command — in either CLI:
--surface/-s and --browser-tier standard|premium work as on task up. --agent/-a is file.py[:fn] or module:fn in Python (fn defaults to agent); in TypeScript it is a
.js/.mjs/.cjs module (or .ts with tsx installed next to it), file.mjs[#export] — the export defaults to
the module’s default export (falling back to a named agent export), and #name picks a specific named export.
Both are called as fn(task) where task is a small TaskRun object: .runId/.run_id, .instruction, .model,
.sessions (alias -> WorldSession), .manifest (the same object task up prints), .api(alias) -> (url, token) (a tuple in Python, {url, token} in TypeScript), .toolkit() (every world’s tools merged into one, each
name prefixed <alias>.<tool> — a world with none contributes nothing), and .record({...}) to log notes into the
run’s run.json. The agent’s return value is the report graders see: a string, or an object/dict with a report
key (other keys land in run.json as-is). A coroutine (Python) or a Promise (TypeScript) the agent returns is
awaited for you.
Every world is graded with that report (WorldSession.grade’s report argument, sent as {"report": ...} — a
rubric judge reads it as the agent’s account of what it did), one rollout is filed the way worlds run files a
single-world run (versionId anchored to the task’s first world; every world’s own version travels in
metadata.worlds), and every session closes — all of this happens
even when the agent raises, with the error recorded in run.json["agentError"] and the filed evaluation marked
FAILED. --no-file skips filing, --no-trace skips tracing the agent’s own session, --keep-up leaves the
sessions open instead of closing them, --surface/-s asks every world for more than tools (comma-separated or
repeated). The filed run is one line on the Runs tab of every scenario set that holds the task.
A run that fails is still recorded: when a world never comes up or a world’s grade cannot be read, run.json says
why (runError, or gradeError naming the world), every other world’s grade is kept, and the evaluation is filed
FAILED with that reason and no reward; a filing the platform refuses is fileError (exit 1). Ctrl-C or SIGTERM
closes every session the run opened (unless --keep-up), even while the platform is still opening them (the run
waits for its answer). Before the grade, it writes run.json with runError: "interrupted by SIGTERM" (or
SIGINT) and files nothing; after the grade, the filing in flight finishes first. Either way it exits 130 (Ctrl-C)
or 143 (SIGTERM).
The run writes <out-dir>/<run id>/manifest.json, grades.json, run.json and (when the agent gave one)
report.md, where <run id> is wtr-<12 hex>. The agent’s own LLM calls and tool spans are traced under one
platform session named by the run id, with the host and keys the command uses (--host, GATEWAY_*, or the saved
gateway auth login); run.json["traced"] says whether tracing actually ran — it never fails the run. A
GATEWAY_OTLP_ENDPOINT on another host is followed only with the environment’s own GATEWAY_PUBLIC_KEY /
GATEWAY_SECRET_KEY: a saved login’s keys go to their own host only. When tracing cannot start, the run continues
untraced and the process prints one warning saying why (never a key); --no-trace runs untraced without it. The
TypeScript CLI traces through optional @opentelemetry/* packages that npm install -g @withgateway/sdk does not
install: the warning prints the npm install -g line that adds them. See
Group traces into a session for how the TypeScript
SDK does the grouping.
A world whose tools the platform could not read (its compiled tool contract is over the publication cap; the
TypeScript session open and task up print it as warning (tool-contract-unread)) is refused before the agent
starts: task run in both CLIs and the TypeScript worlds run print one line naming the world and the reason, exit 1,
close every session (even with --keep-up) and file nothing. task run-set refuses that task the same way, runs the
set’s other tasks and closes the set run failed. The Python worlds run loads the world in-process and reads its tools
there.
Run every task of a set with one model
worlds task run-set runs each stored task of a scenario set the way task run does, and files them all under one
set run: the set page’s Runs tab then shows who ran it, from which branch, commit and pull request, with which
model, how many tasks are done, the score, and every task’s checks.
- One model per set run. Run it again with another
-mto compare models side by side. - The score is the mean of the tasks that were graded; a check or task with no score is shown as such, never as 0.
- The checkout is read from the working directory (branch, commit, remote without credentials). The pull request
comes from
GATEWAY_PR_URL(andGATEWAY_PR_TITLE), or from GitHub Actions’GITHUB_REFon apull_requestrun. - The run closes as
failed(exit 1) when an agent raised or a task could not run or be graded; a low score is stillcompleted. - Ctrl-C or SIGTERM starts no further task: the task in flight closes its sessions, the run closes as
failednaming it, and the command exits 130 or 143. - From code:
runSet(setId, agent, { model })(TypeScript) orrun_set(set_id, agent, model=...)(Python,from gatewaysdk import run_set).
Edit on the platform: the sandbox
worlds sandbox edits a stored world server-side: open a READY version, write and run there, check the tree against the push gate, commit a new version. Nothing lands on this disk. bench sandbox is the same set of verbs.
When the branch moved while you worked
A commit lands on the version the sandbox opened on. If someone pushed to that branch since, a plaincommit is refused with base_moved and nothing is written. Commit again with --rebase to replay your edits onto the branch’s newest version:
1, writes nothing, and prints one line per file on stderr — conflict: <path> (<reason>), where the reason is both_modified, modify_delete, binary or too_large. Make those files agree with the branch, then commit with --rebase again.
When a commit migrates the world’s data
A contract change that only adds to the stored entities (a new entity, or a field that is not required) carries the world’s imported rows to the new contract. That runs as a job:commit prints the batch, follows it with one line per phase on stderr, and then prints the new version. If the job fails, commit exits 1 and names why; nothing is committed and the tree is as you left it. With --wait 0 it returns at once; follow the batch with gateway worlds data status <batch> --world <slug>. While the batch runs, another commit or import on the sandbox is refused (sandbox_commit_pending, naming the batch). Once it lands, a sandbox used in the last 15 minutes takes the new version right away, unless you changed a file the migration rewrites since the commit; otherwise its next commit, import or exec takes it and tells you what it found (close takes it without a word). sandbox check shows that version and the batch without changing anything.
A --rebase merges from your sandbox’s base when the branch was built from it, or when any edit session that carried that base was closed into main, including one that landed it again or built on it. When the newest such session was discarded, it merges from the last version both share. When that session is still open, still closing, or stopped in conflict, the commit is refused (rebase_unrelated) and names the session and what to do: close an open one (gateway worlds edit close <slug> --session <id>), wait for one that is closing, and close one in conflict over main (--force) or discard it (gateway worlds edit discard <slug> --session <id>); then commit with --rebase again.
Closing keeps your work safe
sandbox close refuses while the tree holds edits you have not committed. It exits 1, closes nothing, and names up to ten of the files:
--force to throw it away. A sandbox nobody touches for seven days is closed by the platform; status shows the deadline, and any verb other than status or list moves it.
If the sandbox restarts
The machine behind a sandbox can stop. The platform then starts a new one and puts your tree back before anything else runs: from the latest automatic snapshot, or from the version you opened on when there is none. The next command tells you once, exits1 and does nothing else:
sandbox check <id>), redo what is missing, and carry on; the next command runs normally. A snapshot taken before your last commit is not used: you get the committed version instead. If the snapshot cannot be read back, the command answers sandbox_restore_failed (HTTP 503) and changes nothing; run it again in a moment. close closes a restored sandbox that holds nothing new; if it holds uncommitted edits, close answers sandbox_recreated once and then refuses as usual. If a command deletes the world root itself, check, commit and import answer sandbox_tree_missing; close the sandbox and open a new one.
Commit before you leave a sandbox, or run
sandbox check <id> to see what is still pending.Serve and run
With a local
python3 3.12 or newer, serve runs the world here in the foreground from the bundled runtime. Without one it opens a hosted api session instead and prints the session id and URL, which you close with worlds session close.
GET /__session/task returns the task id, prompt, served tools with each tool’s description and input schema, and the verifiers that grade: the root world’s under verifiers, each linked world’s under worlds.<alias>. POST /__session/call takes {"tool", "args"}; a linked world’s tool is <alias>.<tool>. POST /__session/grade scores every world on its own state and keys a dependency’s verifier <alias>/<verifier>. POST /__session/reset and POST /__session/seed restore or replace the state, and POST /__session/close ends the session.
{"reward": 1.0, "rewards": {"escalated_to_slack": 1.0, "state_present": 1.0, "slack/state_present": 1.0}}.
worlds run takes a <world> that is either a directory or a platform slug[@ref]. --agent/-a is file.mjs[#export], the same form task run takes: the module’s default export, or the export #name picks, is async (session, task) => void; .ts works when tsx is installed. Choose the work with --task/-t for one task or --all for the whole set, pin with --version-id, add surfaces with --surface/-s, and bound the wait with --timeout-min. Each task opens one live session and the task’s own grader scores the end state.
Test a world, and gate a pull request
A world ships tests for itself undertests/: *.json HTTP cases shaped {name, request, expect}, and pytest files that receive WORLD_URL and WORLD_API_KEY.
test takes a world directory or a platform slug[@versionId] to run the suite on the platform. With [actions] on_push = ["tests"] in gateway-env.toml, the platform runs the same suite on every push, and gateway bench tests <slug> reads those runs.
worlds ci is the pull-request gate in one command. It pushes to the world the directory tracks (the slug in .gateway/.env-metadata.json, so a world created as acme-f7 stays acme-f7; pyproject’s name is only the fallback for an untracked directory), pushes the tree as it is, exactly as bench push does, and runs the tasks the platform lists for the pushed version. Its branch defaults to GITHUB_HEAD_REF, then GITHUB_REF_NAME, then the directory’s recorded branch, and it writes a table into $GITHUB_STEP_SUMMARY under GitHub Actions. --run-config <name> dispatches a hosted run config against the pushed version instead of running a local --agent; --summary-file also writes the JSON summary to a path; --branch/-b and --message/-M control the push.
Hand an agent the packaged guide
worlds skill [name] prints one of the packaged skills, and --install <repo> writes it into that repository’s .claude/ along with any agent it ships. See the CLI overview for the five names.
--install <repo> with no name installs every packaged skill. --list prints the names and exits, and --force overwrites a file whose content differs.
Environment variables these commands read
Three variables beyond the credentials and runtime ones on the CLI overview change howgateway worlds behaves.
Where to go next
- gateway benchmarks — the same commands in benchmark order: tasks, k rollouts per model, results.
- bench & versions — publishing the bundle, branches, tags and hosted runs.
- Put data in a world — the rows file contract in full.
- Mock any vendor API —
connector.toml, captures, handlers and conformance. - Getting started with worlds — the same commands as one walkthrough.