Skip to main content
Every pull request gets a number. The workflow below opens one live session per scenario, points the agent under test at each session’s URL, grades what the world holds at the end, and fails the check when the mean reward is below your bar. The agent talks to the session instead of the real vendor, so the platform sees every call it makes and grades the state it leaves behind. Nothing about the agent changes except a base URL and a header.
Every recipe here is plain workflow YAML around npm install -g @withgateway/sdk and the gateway command, the same shape Run API worlds in CI uses.

What the workflow needs

The workflow needs three repository secrets. The keys decide which project the worlds are read from, so a key pointed at a staging project keeps pull requests away from production data. The CLI reads all three from the environment, so no gateway auth login step is needed. gateway auth status early in the job prints which host and key are in effect, so a missing secret shows up there instead of as a 401 later.

The workflow

The agent job runs the scenarios

gateway worlds run <world> --all asks the world for its task list, opens one live session per task, calls your agent with that session, and runs the task’s own grader against the end state. Every session in the run is pinned to the same world version, resolved once, so a push mid-suite cannot split the results across two worlds. --surface api makes the session answer over HTTP. Without it the session serves the tool surface only, and session.api() has no URL to return. Your agent module default-exports async (session, task) => void and reads two values off the session:
A .ts agent works when tsx or ts-node is installed next to it; otherwise build to JavaScript and point --agent at the output. The summary prints to stdout as JSON: world, agent, mean, and one results entry per scenario carrying task, sessionId, reward, rewards and an error string when the scenario threw instead of grading. Progress lines go to stderr, so redirecting stdout to a file keeps the report clean.
gateway worlds run exits 1 when any scenario threw, whatever the mean was. The || true in the step above hands that decision to the threshold step instead, so one crashed scenario shows up as a zero in the mean rather than as a job that died before printing anything. Drop the || true when you want a single crash to fail the pull request on its own.

The tests job checks the world

gateway worlds test <path> runs the suite the world ships for itself: the tests/*.json HTTP cases and the tests/test_*.py files described in Test a world. It exits 1 when any case fails. Pointing it at a path rather than a slug tests the tree in the pull request’s checkout, which is the version that has not been published yet. With a python3 of 3.12 or newer on the runner, the CLI serves the world from its bundled runtime and runs everything locally — no push, no version, no session. Install pytest alongside it. Without it the Python half of the suite reports a skip, and a skip is not a pass.
The argument is read as a path first. A bare slug that matches a directory in the working directory tests that directory instead of the platform’s copy, so keep the two distinct or pass ./worlds/<dir> deliberately.

When the pull request changes the world itself

Neither job above is gateway worlds ci. worlds test runs the world’s own assertions; worlds ci pushes the checked-out world, runs every scenario against the version it just pushed, and gates on the mean:
Use it when the pull request edits the world’s contract, handlers or data, so the runs grade the change being reviewed. It writes a summary table into $GITHUB_STEP_SUMMARY, takes --summary-file for the JSON, and defaults its branch to GITHUB_HEAD_REF, so each pull request’s runs are tagged with its own branch. Run API worlds in CI has the longer form of the same job — opening the session, masking the token, and polling until it is ready — plus the merge workflow that publishes on push.

The knobs

Parallelism

gateway worlds run opens its sessions one after another. There are three ways to widen it, in increasing order of effort: An agent that needs a base URL calls openSession(world, { task, surfaces: ["api"] }) in a pool of its own. Run many sessions at once has both shapes, and Run a task suite covers runSessions in full.

Warm sessions

worlds run closes each session as it finishes. A session closed warm goes back to a pool a later open can reuse: gateway worlds session close <id> --keep-warm from a shell, reuseWarm: true on POST /api/public/world-sessions, and the default for runSessions. Reuse is per version and per task.

The threshold

Apply the score gate for worlds run in your workflow, in one of three places. Pick a bar from a baseline run of the agent already on your main branch, not from a guess. A world’s scenarios vary in difficulty, so the mean is only meaningful against the same world version.

Where the results show up

In the job. The eval-report.json artifact holds every scenario’s reward and error. Read it when the gate fails. In the dashboard. Open Worlds and select the world. Its Results, Runs and Evals tabs carry what ran against it, and Overview lists live sessions while the job is still running. The section’s All results tab is the matrix across every world in the project. The Tests tab holds the runs of the world’s own suite, including the ones the platform ran on push.

Where to go next