
Open one and read what it prints
open prints session <id> opening — gateway worlds session status <id> on stderr, so a wait cut short keeps the id. The platform answers as soon as it accepts the session and the World Host opens it right after; open waits for ready up to --timeout seconds (600 by default). idleExpiresAt is when the idle reaper closes the session, and is null while it is pinned.
sessionId addresses the session in every other command. api.url is the world’s own HTTP
API for that session alone — the host name carries the session id, so two sessions of the
same world never share an address.
Surfaces: ask for what your agent actually uses
Every session servestools over the platform’s rendezvous. Ask for more with --surface
(short form -s), repeatable or comma-separated.
Ask only for surfaces the world declares.
browser needs ui.
A World Host session’s browser runs on one of two tiers. Pick one with --browser-tier; left
out, your organization’s default applies (and standard when the organization has none).
standard browser is reclaimed, the platform puts a new one behind the same wsUrl and
token. A browser from the SDK’s browser() reconnects on its own within 90 seconds and reopens the
page the agent was on. A read (text, screenshot, title) or a goto that lost its browser is then
retried once. A click, type, press, back or scroll is not replayed, since it may already have
reached the world: the SDK raises BrowserReconnected so the agent looks at the page before acting
again. What lived only in the lost browser is gone: the open page’s unsaved client-side state and
anything typed but not submitted.
If you drive your own Playwright, call session.repairBrowser() (TypeScript) or
session.repair_browser() (Python) after a dropped connection, then reconnect to the same wsUrl.
Opened with --surface ui,browser, the descriptor adds where the page and the browser answer:
ui.entry carries the session token once; the host trades it for a cookie on the first load.
gateway worlds session browse <sessionId> opens that page in the browser surface and prints
what the agent would read.
Send the token the way your client already sends credentials
A World Host session gatesapi.url on a per-session token, printed beside the URL. Send it
as a header on every request, never in the URL — a query token would outlive the request in
access logs.
header or basic scheme also accept the same token the vendor’s
own way, so a client written against the real service works unchanged:
Point an existing integration at a world by changing two values: the base URL
and the key. No code in your client has to know a world exists.
Every session names the engine that serves it
gateway worlds session status <sessionId> prints the engine serving it.
Every world opens on the World Host. The engine is not a choice you make on open: the host serves a world’s tools and HTTP routes from one session.
A benchmark package (harbor, verifiers) runs in a container, where the runner serves its tools. A schema world links other worlds into its session, and the host opens each linked world beside it.
Drive the session
1
Open it on a scenario
--version-id <id> to pin an exact version instead of the world’s head. Every
run then records the version it faced.2
Put the agent to work
Call the world’s tools, or its HTTP routes at
api.url, or both. Both change the same state,
so a record written through the API is the record a tool call reads back. Tools and routes are
separate surfaces: a route answers over HTTP at api.url, and a call by a route’s name answers
400 {"error": "not_a_tool"}, pointing at api.url. A tool may share a route’s name; the call
runs the tool, and the route stays at its path. A task’s tools lists the tools it may call, and
a task that lists any also closes the world’s routes for that session (403 route_not_allowed).A session answers its requests one at a time, and the sessions of one world version take turns.
Under a burst a request waits its turn, or answers 503 {"error": "busy"} with a Retry-After
header: wait that many seconds and send it again. A tool call, a route, a grade, a seed, a reset, or
a read of the session’s state or evidence that runs past 45 seconds answers
504 {"error": "call_timeout"} naming it; it may still finish, and until it does the session’s next
requests answer 503 {"error": "call_running"} with a Retry-After. If the process running the
world’s own code restarts (it also does when a call has run for more than two minutes), the call in
flight answers 502 {"error": "worker_exited"}; send it again, and the session, with any world it
opened beside it, carries on from every change already answered.The session page links to its Traced session once anything is traced under the session’s
id: the platform’s trace of each episode’s tool calls, and your agent’s own traces when it
traces inside the session with the Gateway SDK (with open_session(...) as s: in Python,
session.withTraceContext(fn) in TypeScript). See
Trace inside a world session.3
Look at what the world holds
calls, every call the session answered — a
vendor-route request (operation, path args, query, body) or a tool called by name (its name and
arguments) — with selectors (every string the request carried, wherever it sat), via (the door
it came in by), status, the rows the answer carried and found (rows came back), because an
investigation that only reads leaves nothing else behind. rows counts what the route serves:
a list at its envelope or results path, or a single route’s one row, bare or at that route’s
own envelope. It is null where nothing declares the rows, a tool’s object answer included. A field named as a credential (a password, token, secret, API key, PIN, card
number, CVC, e-signature), an item of a list so named (passwords), a value inside a field
so named (password.value), or a value shaped like one (a provider key, a bearer, a JWT, a
private key) is kept as [redacted]; every other field is kept as the agent sent it. A request
too large to keep (over 100,000 values or 1 MiB of text) names what it lacks in withheld: a
route call keeps its path and query and withholds its body, a tool call its arguments. A long
session’s log keeps its newest calls, and its first seq says how many are gone. A session brought
back from its checkpoint numbers on from that checkpoint’s calls, so the calls made after it are
not counted (Bring a closed session back). A check grades
whenever the calls it can read decide it; where a withheld field, a null found or a call no
longer kept could change the result, it is ungraded, never guessed.
A grader reads it like an entity: { "entity": "calls", "where": { "operation": "get_records_by_email", "found": true }, "assert": "any", "field": "selectors", "op": "contains", "value": "alex@example.com" } requires “the agent searched this and got rows”,
whatever the vendor’s request shape. gateway worlds session export <id> --requests prints it whole,
a gateway worlds serve session’s GET /__session/state carries it as the platform’s does, and
snapshot and calls are reserved names an entity may not take.via tells a click in the world’s UI from a request the agent sent straight to the API. The
host sets it from where the request arrived, never from a header, so an agent cannot claim
another door:So “the agent opened the tab from the app, not the API” is one check:
{ "entity": "calls", "where": { "operation": "open_workspace_tab", "via": "ui" }, "assert": "exists" }. via names the door, not who knocked: anything that holds the session’s UI
cookie or token and calls the UI origin’s prefix is ui. gateway worlds serve tags the same
way: only the UI port’s prefix counts as ui; the same prefix on the session port is the API
origin, so it is api.via checks read the call log of a world with HTTP routes (a connector.toml), on a World
Host session or under gateway worlds serve. When you create, update or validate a stored task,
each world its via checks read (in where or as field, on calls or a cross-world
verifier’s <alias>.calls) must run on the World Host and have routes. Create and update
answer 400 naming the world; validate answers ok: false with it under problems. A check
that names via over a call log without it is ungraded (the grade says why), never a pass.A grade may carry the agent’s own report (session.grade({ report }) in TypeScript,
session.grade(report=...) in Python, report in the body of POST …/calls with
kind: "grade"). A rubric judge reads it as the agent’s claim beside the evidence — a
finding is credited only where the state or the calls bear it out, and a claim the evidence
contradicts counts against it. Assertions and python graders ignore it.4
Start over, or stop
The rest of the session commands
session export writes the calls the platform brokered: call, grade, seed, reset, state
and close. A U+0000 character in a call’s arguments or result is recorded, and handed to the
world, as U+FFFD (the replacement character). --requests also reads the world’s own request log from the live session and appends
each entry as a line with kind: "request". That log is the calls a grader reads, so it holds
the requests your agent made through the browser or straight to api.url, each with its via.
{seq, kind: "request", tool, args: {method, path, args, query, body, via}, result: {status, rows, found}, error: null, completedAt}, where tool is the operation, or the
tool called by name (its method and path are null). Credential fields read [redacted] in
request lines; call lines keep a tool’s arguments as your agent sent them.
--requests reads the running session with one state call, which the
ledger records like any other, so export before you close. A session that
cannot answer fails the command, and nothing is written.Tasks that need several worlds
Some tasks span worlds: an investigation reads a breach feed, a threat-intel API and a ticketing system, and the agent needs all three up at once. A stored task declares every world it needs in aworlds table, and one command brings them all up.
slug), optionally one of its branches or tags (ref, default main), and that world’s
share of the task: its own seed, grader and tools, the same three fields a single-world
task carries. With worlds present the top-level seed, grader and tools are refused
(they have no world to apply to), and the task is not pinned to any one world. The
instruction and metadata are shared by every world.
task up opens one session per alias in one platform call and prints one manifest:
session open prints for that world. up is all or nothing: if any
world refuses to open, the ones that did are closed and the error names every alias that
failed — nothing is ever half up. up also takes a spec file directly (task up ./orion.json)
for a task you have not stored.
Each ref resolves through your edit view: when you have an open edit session on a world
and it has written a version (see gateway worlds edit status), that world’s main comes up
at your session’s tip; everyone else still gets the shared main until you close the edit.
task create and task validate check the same versions, so what they accept is what up
opens.
"private": true means exactly one thing: your open edit session’s tip stood in for main.
It is false for a tag or any other branch, and for an edit session that has not written
anything yet. Every manifest entry and every validate row carries it; validate on a task
pinned to one world carries it at the top level (false for --version-id), and a
multi-world answer’s top level and an alias that did not resolve say null.
The later commands take either the manifest or the task id. With a task id, status and
grade find the stored task’s live sessions (and refuse when an alias is up twice from two
ups — pass the manifest then); down closes every live session of the task.
task grade answers {reward, rewards, worlds, ungraded, verifiers, errors}: worlds holds
each alias’s own grade, rewards flattens them as <alias> and <alias>.<check>, reward is
the mean of the worlds that answered a number, and a world with no grader (or a rubric judge
still pending) is listed under ungraded rather than counted as 0. A grader that could not
grade is named under errors; the reward is then null and the command exits 1. task export writes every
world’s call ledger as one JSONL file with a world field on each line; --requests adds each
world’s own request log the same way (wrote N calls and M requests from K worlds to <file>).
Under the hood this is POST /api/public/world-task-sessions with {task: {id} | spec, surfaces?, reuseWarm?, browserTier?} (browserTier applies to every world), answering the manifest, and GET /api/public/world-task-sessions?taskId=
listing a task’s live sessions. Open a multi-world task through task up or this endpoint.
Edit several worlds together
Every change you push to a world (bench push, a sandbox commit, a data import) lands on your
edit session for that world: a private copy only your key sees, which becomes the world’s shared
main when you close it, or after 60 minutes without an edit. When one change spans the worlds of
a task, close them together, so main never holds half of it:
statuslists, per alias, your open session (edit count, base, tip) and what closing now would do:fast-forward,rebase(mainmoved on other files) orconflictwith the paths, plus any of your sessionsparkedin conflict there.unsealednames the worlds holding edits only you see: your owntask upopens them at your tip ("private": true), while colleagues, run configs and schedules still get the sharedmainuntil you close them.closeseals every world in one step: everymainmoves, or none does. If any world conflicts, is not in your project, names arefother thanmain, or holds a session of yours parked in conflict, nothing is sealed and the error names every world’s problem (exit 1); a parked session is resolved withgateway worlds edit close <slug> --session <id> --forceoredit discard. A world with no open edit is skipped and listed underskipped;--require-allrefuses instead.--forcekeeps your files wheremainchanged the same paths.- The answer is
{groupId, worlds: {alias: {slug, session}}, skipped}.rollback --group <groupId>returns every world’s source and data to themainthe group replaced (a colleague’s change that landed before yours stays), again in one step, as new sessions you can roll back too. It is refused when anything reached a world’smainafter the group — a later session or a write outside sessions — unless--force. A world whose seal leftmainwhere it was (someone had already sealed the same files) has nothing to undo and is listed underskipped; the other worlds still roll back together. <task>is a task id, a spec file, or the manifesttask up --outwrote. Only your own open sessions are sealed; close before a session idles out (60 minutes), or that world seals on its own.
POST /api/public/world-task-edits/status|close|rollback with {task: {id} | spec, …}; the MCP tools are get_task_edit, close_task_edit and rollback_task_edit. For one
world, gateway worlds edit status|close|rollback <slug> does the same.
Verifiers that span worlds
A world’sgrader only sees that world. A scenario question that spans worlds — was the
company officer found in the registry the owner of the wallet seen on chain? did the agent
consult every source before concluding? — lives on the task instead, as a verifier: the
same assertions or python grader, with a scope of aliases it reads.
<alias>.<entity> and
its call log <alias>.calls — the naming a world’s dependencies already use, so a check or
a verify.py written for one world reads the merged state the same way. The agent’s
report sits at the top, as for any grader. A scope names two or more of the task’s
aliases; a rubric cannot be a verifier (a judge reads one world), and a verifier cannot be
anchored on a world whose own grader is a rubric.
A python grader or verifier prints one flat JSON object as its last line: reward plus one
number per check, such as {"reward": 0.5, "card_moved": 1, "assigned": 0}. Each other key
becomes a check under the world’s alias (asana.card_moved). A check is a number from 0 to 1,
true or false (1 or 0), or null for a check it could not grade; a string reads as ungraded
too. A nested object or array, such as checks nested under rewards, a number outside 0 to 1, or
NaN / Infinity fails the grade with an error that names the key (and the value, for a
number). Print diagnostics such as counts to stderr.
Verifiers run on the platform, never on your machine: each one runs on the World Host
session of the first alias in its scope, and the platform reads the other scoped worlds’
evidence from their own sessions when the grade is called — task grade and the SDKs name
the task’s other sessions on every grade for you. task grade then answers:
x.<name>, so no world may be aliased x or x.<…>:
creating or validating such a task is refused. The task’s reward is the mean over the worlds
and the verifiers that answered a number. A verifier that could not grade (a scoped
world’s session gone, a check naming an entity no scoped world holds, a verify.py that
failed) is named under errors as x.<name>, with the reason also under
verifiers.<name>.error; the task’s reward is then null, never the mean of the rest, and
task run files the run failed with that reason. No world’s contract vets a cross-world
check’s entity, so a misspelt one is an error rather than a silently passing absent.
Grading a session by hand (session grade --world alias=sessionId …, or POST …/calls
with kind: "grade" and worlds: {alias: sessionId}) names the other sessions the same
way. What the caller got wrong is refused with its name: a scoped alias left out, or a
session that belongs to another task than the alias’s world.
Keep a session up past the idle timeout
A session closes after 30 idle minutes unless it is pinned. Traffic counts as activity, including plain HTTP toapi.url, so a session in use stays up on its own.
Pin a session when its URL has to answer around the clock — a demo, a long-lived integration
test, a world someone else is pointing a client at. Four ways do the same thing:
close is the only way out.
Bring a closed session back
A session that closed — by you or by the idle reaper — can come back under the same id:status prints. The platform takes that checkpoint about once a minute while the session runs.
Work done after that export is not in it. The checkpoint holds the world’s request log as well as its
rows, so graders and session export --requests still see the requests made up to the checkpoint.
Requests made after it are gone with the rows they wrote, and the log numbers on from the
checkpoint’s last request, so nothing marks them. A session recovered after its host was lost
resumes the same way.
The command exits 1 while the
session still runs (session_not_closed), when it was never checkpointed
(session_no_checkpoint — open a new session instead), while another reopen is under way
(session_reopen_in_progress — retry), or when the world version it ran on is no longer available
(session_version_unavailable — open a new session). Closing a session the platform is bringing
back after its host was lost answers session_reopen_in_progress too: retry once it runs.
A session closed within about a minute of opening has no checkpoint yet and answers
session_no_checkpoint.
Verify
Three checks, cheapest first. Ask the platform.gateway worlds session status <sessionId> prints state, engine,
tools and surfaces. ready means a call will be answered, not that the job was accepted.
Ask the world. A request to api.url with the token returns the vendor’s own response;
without the token it returns 401, and after the session closes it returns 404.
Close the sessions you open. A leaked session holds its container until the
idle reaper finds it, and a pinned leaked session holds it forever.
--keep-warm on close gives the container back to the warm pool so the next
session for the same task starts warm.Where to go next
- Run many sessions at once for suites, concurrency and CI.
- Simulations for your users to serve a world per customer.
- Put data in a world for seeding a live session with rows.
- Worlds client to open and drive sessions from Python.