tests/. They run against the version they ship with — locally while you author, on the platform on demand, and on every push when the manifest asks for it.

A test asserts the world; an eval grades the agent
Both produce a pass rate. They answer different questions.
An eval that drops from 0.9 to 0.4 is ambiguous until the tests are green. With green tests, the drop belongs to the agent.
A world with no
tests/ directory passes with zero tests. Nothing is asserted, so nothing fails and nothing is proven. The exception is a world whose gateway-env.toml declares [actions] on_push = ["tests"] and ships no case: it promised a suite, so a missing or empty tests/ fails.What a world ships
Three things, all inside the world directory:test_*.py or *_test.py are collected as pytest. Every .json file under tests/ is read as cases, so subdirectories work. tasks.json and scenarios/ are not tests; see the two bundled task formats.
Add the action to gateway-env.toml at the world root to run the suite on every push:
Write an HTTP case
A case file is{"cases": [...]}, or a bare list. Each case is a request and what the answer must contain.
name is what the report prints. group and description are optional and travel into the report unchanged.
The request half
With
auth: true the runner authenticates through whatever scheme connector.toml declares — header, bearer, basic, query, or an oauth2 world through its own client-credentials grant. Set auth: false to assert the unauthenticated path; it is the only way to test the vendor’s 401 body.
The expect half
json is a subset check, not equality, so a case survives the vendor adding a field. Assert the fields your agent actually reads and leave the rest alone.
Write a tool case
A case may call a tool the way an agent does instead of a route.tool names a root tool, or a dependency’s as alias.tool; args are its arguments. The answer checked is the tool’s {status, body}. A refused call (unknown tool, arguments off the schema) answers 400 with {error, detail}, and a case may expect it.
{"scenario": "default", "cases": [...]}; the session opens seeded with it. Two files naming different scenarios are refused, because one session runs them all.
Write pytest for a sequence
A JSON case asserts one request. A sequence — create, then poll, then read back — needs Python. One session serves the whole suite: the JSON files in name order, then the pytest files, all withWORLD_URL and WORLD_API_KEY in the environment. A pytest file sees what the cases before it wrote, so count rows by id, not by total.
Start from a clean world
POST {WORLD_URL}/__session/reset returns the session to the world’s initial state; on a world with connector.toml, POST /__world/reset with the header X-World-Admin-Token: $WORLD_ADMIN_TOKEN (set for every Python test) does the same. A reset discards what the JSON cases wrote and does not re-apply the scenario a case file named. A test that needs the scenario’s rows back seeds them again: POST /__session/seed with {"rows": {"<entity>": [...]}, "mode": "replace"}.
Python tests fail, never skip, when pytest is missing. The runner uses the first interpreter that has pytest: the one running the suite, then
$VIRTUAL_ENV, a .venv beside the world or in the working directory, then python3 on PATH. With none, the report names the interpreter and the install command, the python tests line prints that reason, and gateway worlds test exits 1. The platform ships pytest.Run them
1
While you author, against the tree on disk
--scenario <name> to start the world in one of its scenarios instead of the base state.2
On the platform, against a published version
main for your key is the version you pushed, so that is the one tested; other keys keep testing the shared main. The command starts the run and polls until it settles, printing the same report. Its last line ends with (run <id>, trigger <trigger>, on <versionId>), and --json and --out carry trigger too.3
On every push, without asking
With
[actions] on_push = ["tests"] in gateway-env.toml, each pushed version gets its suite run against it. The run is fire-and-forget: a push never waits on its tests and never fails because of them. The result lands on the version for the Tests tab and gateway bench tests to read.Read the report
The report groups cases under theirgroup and prints each description beneath the name.
- The world’s Tests tab groups a run by
groupwith a passed/total count per group, and offers Run tests now. gateway bench tests <slug>lists the runs the platform kept, newest first;--run <id>prints one report in full.POST /api/public/worlds/{slug}/testsruns the suite over plain HTTP. With{"wait": false}it answers immediately with a queued run, andGET /api/public/worlds/{slug}/tests/{runId}reads it until it settles.
push, manual, or api), the version it ran against, and the full report.
Two things that catch people out
The on-push run sees the pushed tree, and only the pushed tree
A world whose rows arrive on a later data version has no evidence in the tree at the moment of a code push. A JSON case that assertslength: {path: "data", eq: 42} then fails for a reason unrelated to correctness.
Keep JSON cases bare-tree-safe — status codes, envelope shape, error bodies, auth behaviour — and put count assertions in pytest that skips itself when the evidence is absent:
Under the runner, the world’s credential is pinned to conform
The test runner serves the world with its API key pinned to the literal string conform, then presents that credential on every case with auth: true. Bearer, header, basic and query schemes all use the pin; an oauth2 world is taken through its own client-credentials grant automatically.
For a password grant, write the token request as its own case with auth: false, using conform as the password:
Tasks are not tests: the two bundled formats
A test asserts the world. A bundled task is what a session is graded on. Two formats exist, read by different engines.
A session task supplied at open (
--task-file) needs neither file: the World Host grades its grader, and the tools engine grades it with every verifier the world ships. See a world per tenant, a task per user.
Where to go next
- Run API worlds in CI to put the suite in a GitHub Actions workflow beside an agent smoke run.
- What makes a world good for where the cases come from: one described case per route the client actually calls.
- Mock any vendor API for
connector.toml, capture, handlers and conformance against the vendor’s OpenAPI. - Getting started with worlds for the whole loop from empty directory to graded run.