Skip to main content
A world ships its own tests under tests/. They run against the version they ship with — locally while you author, on the platform on demand, and on every push when the manifest asks for it. A world's Tests tab: each run with its cases grouped and described

A test asserts the world; an eval grades the agent

Both produce a pass rate. They answer different questions. An eval that drops from 0.9 to 0.4 is ambiguous until the tests are green. With green tests, the drop belongs to the agent.
A world with no tests/ directory passes with zero tests. Nothing is asserted, so nothing fails and nothing is proven. The exception is a world whose gateway-env.toml declares [actions] on_push = ["tests"] and ships no case: it promised a suite, so a missing or empty tests/ fails.

What a world ships

Three things, all inside the world directory:
Files named test_*.py or *_test.py are collected as pytest. Every .json file under tests/ is read as cases, so subdirectories work. tasks.json and scenarios/ are not tests; see the two bundled task formats. Add the action to gateway-env.toml at the world root to run the suite on every push:

Write an HTTP case

A case file is {"cases": [...]}, or a bare list. Each case is a request and what the answer must contain.
name is what the report prints. group and description are optional and travel into the report unchanged.

The request half

With auth: true the runner authenticates through whatever scheme connector.toml declares — header, bearer, basic, query, or an oauth2 world through its own client-credentials grant. Set auth: false to assert the unauthenticated path; it is the only way to test the vendor’s 401 body.

The expect half

json is a subset check, not equality, so a case survives the vendor adding a field. Assert the fields your agent actually reads and leave the rest alone.

Write a tool case

A case may call a tool the way an agent does instead of a route. tool names a root tool, or a dependency’s as alias.tool; args are its arguments. The answer checked is the tool’s {status, body}. A refused call (unknown tool, arguments off the schema) answers 400 with {error, detail}, and a case may expect it.
A case file may name the scenario its cases assume, {"scenario": "default", "cases": [...]}; the session opens seeded with it. Two files naming different scenarios are refused, because one session runs them all.

Write pytest for a sequence

A JSON case asserts one request. A sequence — create, then poll, then read back — needs Python. One session serves the whole suite: the JSON files in name order, then the pytest files, all with WORLD_URL and WORLD_API_KEY in the environment. A pytest file sees what the cases before it wrote, so count rows by id, not by total.

Start from a clean world

POST {WORLD_URL}/__session/reset returns the session to the world’s initial state; on a world with connector.toml, POST /__world/reset with the header X-World-Admin-Token: $WORLD_ADMIN_TOKEN (set for every Python test) does the same. A reset discards what the JSON cases wrote and does not re-apply the scenario a case file named. A test that needs the scenario’s rows back seeds them again: POST /__session/seed with {"rows": {"<entity>": [...]}, "mode": "replace"}.
Python tests fail, never skip, when pytest is missing. The runner uses the first interpreter that has pytest: the one running the suite, then $VIRTUAL_ENV, a .venv beside the world or in the working directory, then python3 on PATH. With none, the report names the interpreter and the install command, the python tests line prints that reason, and gateway worlds test exits 1. The platform ships pytest.

Run them

1

While you author, against the tree on disk

The world is served in-process and every case runs against it. Exit code 1 on any failure, so the same command works as a pre-commit hook. Add --scenario <name> to start the world in one of its scenarios instead of the base state.
2

On the platform, against a published version

With a slug rather than a path, the platform runs the suite the published version ships. Right after you push, while your edit is not yet sealed, main for your key is the version you pushed, so that is the one tested; other keys keep testing the shared main. The command starts the run and polls until it settles, printing the same report. Its last line ends with (run <id>, trigger <trigger>, on <versionId>), and --json and --out carry trigger too.
3

On every push, without asking

With [actions] on_push = ["tests"] in gateway-env.toml, each pushed version gets its suite run against it. The run is fire-and-forget: a push never waits on its tests and never fails because of them. The result lands on the version for the Tests tab and gateway bench tests to read.

Read the report

The report groups cases under their group and prints each description beneath the name.
Three other ways to the same runs:
  • The world’s Tests tab groups a run by group with a passed/total count per group, and offers Run tests now.
  • gateway bench tests <slug> lists the runs the platform kept, newest first; --run <id> prints one report in full.
  • POST /api/public/worlds/{slug}/tests runs the suite over plain HTTP. With {"wait": false} it answers immediately with a queued run, and GET /api/public/worlds/{slug}/tests/{runId} reads it until it settles.
Every run records its trigger (push, manual, or api), the version it ran against, and the full report.

Two things that catch people out

The on-push run sees the pushed tree, and only the pushed tree

A world whose rows arrive on a later data version has no evidence in the tree at the moment of a code push. A JSON case that asserts length: {path: "data", eq: 42} then fails for a reason unrelated to correctness. Keep JSON cases bare-tree-safe — status codes, envelope shape, error bodies, auth behaviour — and put count assertions in pytest that skips itself when the evidence is absent:

Under the runner, the world’s credential is pinned to conform

The test runner serves the world with its API key pinned to the literal string conform, then presents that credential on every case with auth: true. Bearer, header, basic and query schemes all use the pin; an oauth2 world is taken through its own client-credentials grant automatically. For a password grant, write the token request as its own case with auth: false, using conform as the password:

Tasks are not tests: the two bundled formats

A test asserts the world. A bundled task is what a session is graded on. Two formats exist, read by different engines. A session task supplied at open (--task-file) needs neither file: the World Host grades its grader, and the tools engine grades it with every verifier the world ships. See a world per tenant, a task per user.

Where to go next