> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Test a world

> A world ships its own tests. Write HTTP cases, tool cases and pytest files under tests/, run them locally, on the platform, or on every push, and read the report grouped by route.

A world ships its own tests under `tests/`. They run against the version they ship with — locally while you author, on the platform on demand, and on every push when the manifest asks for it.

<img src="https://mintcdn.com/surface-d3d890e1/I9MKHQA4beHtjYQ2/screenshots/world-tests-tab.png?fit=max&auto=format&n=I9MKHQA4beHtjYQ2&q=85&s=353cdd1ce9ed229e8c5c837ab6c958d2" alt="A world's Tests tab: each run with its cases grouped and described" width="2880" height="1800" data-path="screenshots/world-tests-tab.png" />

## A test asserts the world; an eval grades the agent

Both produce a pass rate. They answer different questions.

| | What it checks | When it fails |
| - | - | - |
| **Test** | The world answers the way its author says it must | The world is wrong |
| **[Eval](/glossary#eval)** | The agent did the task well enough | The agent is wrong |

An eval that drops from 0.9 to 0.4 is ambiguous until the tests are green. With green tests, the drop belongs to the agent.

<Info>
  A world with no `tests/` directory passes with zero tests. Nothing is asserted, so nothing fails and nothing is proven. The exception is a world whose `gateway-env.toml` declares `[actions] on_push = ["tests"]` and ships no case: it promised a suite, so a missing or empty `tests/` fails.
</Info>

## What a world ships

Three things, all inside the world directory:

```
<world>/tests/*.json        HTTP cases against the served vendor API, and tool cases
<world>/tests/test_*.py     pytest, with WORLD_URL and WORLD_API_KEY in the environment
gateway-env.toml            [actions] on_push = ["tests"]
```

Files named `test_*.py` or `*_test.py` are collected as pytest. Every `.json` file under `tests/` is read as cases, so subdirectories work. `tasks.json` and `scenarios/` are not tests; see [the two bundled task formats](#tasks-are-not-tests-the-two-bundled-formats).

Add the action to `gateway-env.toml` at the world root to run the suite on every push:

```toml theme={null}
[actions]
on_push = ["tests"]
```

## Write an HTTP case

A case file is `{"cases": [...]}`, or a bare list. Each case is a request and what the answer must contain.

```json theme={null}
{
  "cases": [
    {
      "name": "an address with no activity answers 200 and an empty list",
      "group": "address activities",
      "description": "The vendor documents no 404 on this route — an unseen address is an empty result, not an error.",
      "request": {
        "method": "GET",
        "path": "/api/v1/addresses/0xabc/activities",
        "auth": true
      },
      "expect": { "status": 200, "json": [] }
    }
  ]
}
```

`name` is what the report prints. `group` and `description` are optional and travel into the report unchanged.

### The request half

| Field | Default | What it does |
| - | - | - |
| `method` | `GET` | HTTP method |
| `path` | required | The vendor path; the base URL's prefix may be left off |
| `query` | none | Query parameters as an object |
| `headers` | none | Extra request headers |
| `body` | none | JSON body, sent with `Content-Type: application/json` |
| `auth` | `true` | Present the world's own credential the way a real client would |

With `auth: true` the runner authenticates through whatever scheme `connector.toml` declares — `header`, `bearer`, `basic`, `query`, or an `oauth2` world through its own client-credentials grant. Set `auth: false` to assert the unauthenticated path; it is the only way to test the vendor's 401 body.

### The expect half

| Field | What it asserts |
| - | - |
| `status` | An integer, or a list of acceptable integers |
| `json` | A subset the answer must contain: object keys recursively, list items by index |
| `length` | `{path, eq \| min \| max}` over a dotted path into the answer; `$` is the root |
| `contains` | Substrings that must appear in the raw body |
| `headers` | A subset of the response headers |

`json` is a subset check, not equality, so a case survives the vendor adding a field. Assert the fields your agent actually reads and leave the rest alone.

```json theme={null}
{
  "name": "the ledger lists a settled charge with its amount in minor units",
  "group": "charges",
  "request": { "method": "GET", "path": "/v1/charges", "query": { "status": "settled" } },
  "expect": {
    "status": 200,
    "length": { "path": "data", "min": 1 },
    "json": { "data": [{ "currency": "usd", "status": "settled" }] }
  }
}
```

## Write a tool case

A case may call a tool the way an agent does instead of a route. `tool` names a root tool, or a dependency's as `alias.tool`; `args` are its arguments. The answer checked is the tool's `{status, body}`. A refused call (unknown tool, arguments off the schema) answers 400 with `{error, detail}`, and a case may expect it.

```json theme={null}
{
  "name": "send_message refuses an empty body",
  "tool": "send_message",
  "args": { "conversation_id": "c1", "body": "" },
  "expect": { "status": 400, "contains": ["body must not be empty"] }
}
```

A case file may name the scenario its cases assume, `{"scenario": "default", "cases": [...]}`; the session opens seeded with it. Two files naming different scenarios are refused, because one session runs them all.

## Write pytest for a sequence

A JSON case asserts one request. A sequence — create, then poll, then read back — needs Python. One session serves the whole suite: the JSON files in name order, then the pytest files, all with `WORLD_URL` and `WORLD_API_KEY` in the environment. A pytest file sees what the cases before it wrote, so count rows by id, not by total.

```python theme={null}
import os
import urllib.request
import json

URL = os.environ["WORLD_URL"]
KEY = os.environ["WORLD_API_KEY"]


def call(method, path, body=None):
    data = json.dumps(body).encode() if body is not None else None
    request = urllib.request.Request(
        URL + path,
        method=method,
        data=data,
        headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"},
    )
    with urllib.request.urlopen(request) as reply:
        return reply.status, json.loads(reply.read() or b"null")


def test_a_created_search_is_readable_by_its_id():
    status, created = call("POST", "/v1/searches", {"name": "probe"})
    assert status == 201
    status, found = call("GET", f"/v1/searches/{created['id']}")
    assert status == 200
    assert found["name"] == "probe"
```

### Start from a clean world

`POST {WORLD_URL}/__session/reset` returns the session to the world's initial state; on a world with `connector.toml`, `POST /__world/reset` with the header `X-World-Admin-Token: $WORLD_ADMIN_TOKEN` (set for every Python test) does the same. A reset discards what the JSON cases wrote and does not re-apply the scenario a case file named. A test that needs the scenario's rows back seeds them again: `POST /__session/seed` with `{"rows": {"<entity>": [...]}, "mode": "replace"}`.

<Info>
  Python tests **fail, never skip**, when pytest is missing. The runner uses the first interpreter that has pytest: the one running the suite, then `$VIRTUAL_ENV`, a `.venv` beside the world or in the working directory, then `python3` on `PATH`. With none, the report names the interpreter and the install command, the `python tests` line prints that reason, and `gateway worlds test` exits 1. The platform ships pytest.
</Info>

## Run them

<Steps>
  <Step title="While you author, against the tree on disk">
    ```bash theme={null}
    gateway worlds test ./vendor-world
    ```

    The world is served in-process and every case runs against it. Exit code 1 on any failure, so the same command works as a pre-commit hook. Add `--scenario <name>` to start the world in one of its scenarios instead of the base state.
  </Step>

  <Step title="On the platform, against a published version">
    ```bash theme={null}
    gateway worlds test vendor-world              # main, as your key sees it
    gateway worlds test vendor-world@<versionId>  # a pinned version
    ```

    With a slug rather than a path, the platform runs the suite the published version ships. Right after you push, while your edit is not yet sealed, `main` for your key is the version you pushed, so that is the one tested; other keys keep testing the shared `main`. The command starts the run and polls until it settles, printing the same report. Its last line ends with `(run <id>, trigger <trigger>, on <versionId>)`, and `--json` and `--out` carry `trigger` too.
  </Step>

  <Step title="On every push, without asking">
    With `[actions] on_push = ["tests"]` in `gateway-env.toml`, each pushed version gets its suite run against it. The run is fire-and-forget: a push never waits on its tests and never fails because of them. The result lands on the version for the Tests tab and `gateway bench tests` to read.
  </Step>
</Steps>

## Read the report

The report groups cases under their `group` and prints each `description` beneath the name.

```
== address activities
 ok   an address with no activity answers 200 and an empty list  GET /api/v1/addresses/0xabc/activities
        The vendor documents no 404 on this route — an unseen address is an empty result, not an error.
FAIL  a screened address carries its risk score  GET /api/v1/addresses/0xdef
        ! $.riskScore: missing
 ok   python tests (tests/test_searches.py)  3 tests, 0 failures, 0 errors
passed: 4 passed, 0 failed, 0 skipped
```

Three other ways to the same runs:

* **The world's Tests tab** groups a run by `group` with a passed/total count per group, and offers **Run tests now**.
* `gateway bench tests <slug>` lists the runs the platform kept, newest first; `--run <id>` prints one report in full.
* `POST /api/public/worlds/{slug}/tests` runs the suite over plain HTTP. With `{"wait": false}` it answers immediately with a queued run, and `GET /api/public/worlds/{slug}/tests/{runId}` reads it until it settles.

Every run records its trigger (`push`, `manual`, or `api`), the version it ran against, and the full report.

## Two things that catch people out

### The on-push run sees the pushed tree, and only the pushed tree

A world whose rows arrive on a later data version has no evidence in the tree at the moment of a code push. A JSON case that asserts `length: {path: "data", eq: 42}` then fails for a reason unrelated to correctness.

Keep JSON cases **bare-tree-safe** — status codes, envelope shape, error bodies, auth behaviour — and put count assertions in pytest that skips itself when the evidence is absent:

```python theme={null}
import os
import pytest

def test_every_sanctioned_entity_is_screenable():
    rows = fetch("/v1/entities?list=sanctions")
    if not rows:
        pytest.skip("no sanctions evidence on this version — counts are asserted after a data publish")
    assert len(rows) == 412
```

### Under the runner, the world's credential is pinned to `conform`

The test runner serves the world with its API key pinned to the literal string `conform`, then presents that credential on every case with `auth: true`. Bearer, header, basic and query schemes all use the pin; an `oauth2` world is taken through its own client-credentials grant automatically.

For a **password grant**, write the token request as its own case with `auth: false`, using `conform` as the password:

```json theme={null}
{
  "name": "the password grant mints a bearer",
  "group": "auth",
  "request": {
    "method": "POST",
    "path": "/oauth/token",
    "auth": false,
    "body": { "grant_type": "password", "email": "ops@acme.example", "password": "conform" }
  },
  "expect": { "status": 200, "json": { "token_type": "Bearer" } }
}
```

## Tasks are not tests: the two bundled formats

A test asserts the world. A bundled task is what a session is graded on. Two formats exist, read by different engines.

| Format | Files | Read by |
| - | - | - |
| Scenario task | `scenarios/<name>/scenario.toml`, `seed/*.json` overlays, `instruction.md`, `checks/verify.py` | The World Host: `verify.py` grades over the state dump. Every engine reads the `seed/` overlays to start a session in the scenario. |
| Verifier task | `tasks.json` (or `[[tasks]]` in `gateway-env.toml`) + `verifiers/*.sql` | The tools engine: `gateway worlds test`, `call` and `serve` on a directory, and hosted sessions of a world stored as a tree (`gateway worlds create --from <dir>`), including a world with `[dependencies]`. Each task names the verifiers that grade it; with no `tasks.json`, every verifier is a task. |

A session task supplied at open (`--task-file`) needs neither file: the World Host grades its `grader`, and the tools engine grades it with every verifier the world ships. See [a world per tenant, a task per user](/worlds/for-your-users).

## Where to go next

* [Run API worlds in CI](/worlds/ci) to put the suite in a GitHub Actions workflow beside an agent smoke run.
* [What makes a world good](/worlds/good-world) for where the cases come from: one described case per route the client actually calls.
* [Mock any vendor API](/worlds/custom-connections) for `connector.toml`, capture, handlers and conformance against the vendor's OpenAPI.
* [Getting started with worlds](/worlds/getting-started) for the whole loop from empty directory to graded run.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.