Every recipe here is plain workflow YAML around
npm install -g @withgateway/sdk and the gateway command, the same shape
Run API worlds in CI uses.What the workflow needs
The workflow needs three repository secrets. The keys decide which project the worlds are read from, so a key pointed at a staging project keeps pull requests away from production data.
The CLI reads all three from the environment, so no
gateway auth login step is needed. gateway auth status early in the job prints which host and key are in effect, so a missing secret shows up there instead of as a 401 later.
The workflow
The agent job runs the scenarios
gateway worlds run <world> --all asks the world for its task list, opens one live session per task, calls your agent with that session, and runs the task’s own grader against the end state. Every session in the run is pinned to the same world version, resolved once, so a push mid-suite cannot split the results across two worlds.
--surface api makes the session answer over HTTP. Without it the session serves the tool surface only, and session.api() has no URL to return.
Your agent module default-exports async (session, task) => void and reads two values off the session:
.ts agent works when tsx or ts-node is installed next to it; otherwise build to JavaScript and point --agent at the output.
The summary prints to stdout as JSON: world, agent, mean, and one results entry per scenario carrying task, sessionId, reward, rewards and an error string when the scenario threw instead of grading. Progress lines go to stderr, so redirecting stdout to a file keeps the report clean.
gateway worlds run exits 1 when any scenario threw, whatever the mean
was. The || true in the step above hands that decision to the threshold
step instead, so one crashed scenario shows up as a zero in the mean rather
than as a job that died before printing anything. Drop the || true when you
want a single crash to fail the pull request on its own.The tests job checks the world
gateway worlds test <path> runs the suite the world ships for itself: the tests/*.json HTTP cases and the tests/test_*.py files described in Test a world. It exits 1 when any case fails.
Pointing it at a path rather than a slug tests the tree in the pull request’s checkout, which is the version that has not been published yet. With a python3 of 3.12 or newer on the runner, the CLI serves the world from its bundled runtime and runs everything locally — no push, no version, no session.
Install pytest alongside it. Without it the Python half of the suite reports a skip, and a skip is not a pass.
The argument is read as a path first. A bare slug that matches a directory in
the working directory tests that directory instead of the platform’s copy, so
keep the two distinct or pass
./worlds/<dir> deliberately.When the pull request changes the world itself
Neither job above isgateway worlds ci. worlds test runs the world’s own assertions; worlds ci pushes the checked-out world, runs every scenario against the version it just pushed, and gates on the mean:
$GITHUB_STEP_SUMMARY, takes --summary-file for the JSON, and defaults its branch to GITHUB_HEAD_REF, so each pull request’s runs are tagged with its own branch.
Run API worlds in CI has the longer form of the same job — opening the session, masking the token, and polling until it is ready — plus the merge workflow that publishes on push.
The knobs
Parallelism
gateway worlds run opens its sessions one after another. There are three ways to widen it, in increasing order of effort:
An agent that needs a base URL calls
openSession(world, { task, surfaces: ["api"] }) in a pool of its own. Run many sessions at once has both shapes, and Run a task suite covers runSessions in full.
Warm sessions
worlds run closes each session as it finishes. A session closed warm goes back to a pool a later open can reuse: gateway worlds session close <id> --keep-warm from a shell, reuseWarm: true on POST /api/public/world-sessions, and the default for runSessions. Reuse is per version and per task.
The threshold
Apply the score gate forworlds run in your workflow, in one of three places.
Pick a bar from a baseline run of the agent already on your main branch, not from a guess. A world’s scenarios vary in difficulty, so the mean is only meaningful against the same world version.
Where the results show up
In the job. Theeval-report.json artifact holds every scenario’s reward and error. Read it when the gate fails.
In the dashboard. Open Worlds and select the world. Its Results, Runs and Evals tabs carry what ran against it, and Overview lists live sessions while the job is still running. The section’s All results tab is the matrix across every world in the project. The Tests tab holds the runs of the world’s own suite, including the ones the platform ran on push.
Where to go next
- Run API worlds in CI for the hand-rolled session job and the publish-on-merge workflow.
- Test a world for the case format and what the report says.
- Run many sessions at once for the parallel shapes.
- Simulations from real data for the other recipe.