> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# gateway benchmarks

> Build, validate, dry-run, push, run and read a benchmark from the terminal. A benchmark is a world with tasks or a container benchmark (harbor, verifiers); one set of commands, the kind read from the tree.

A benchmark is either a [world](/worlds) with `[[tasks]]` (a prompt and the SQL verifiers that grade the world's end state; run on the World Host) or a **container benchmark** (a harbor task tree graded by its own `tests/test.sh` inside a container, or a verifiers package graded by Python reward functions; run by the benchmark runner). `gateway benchmarks` is one set of commands for both. The kind is read from the directory, never guessed.

| Command | World tree | Harbor tree / verifiers package |
| - | - | - |
| `benchmarks init <dir> [--kind …]` | `worlds schema init` + `worlds schema task add`, plus `benchmark.json` | `--kind harbor`: one task with instruction, image, grader and solution; `--kind verifiers`: a Taskset package |
| `benchmarks validate <dir>` | `worlds schema check`; exit 2 without a task | every task's files, the reward contract in `tests/test.sh`, and the kind the runner will read |
| `benchmarks dry-run <dir> [--task …]` | open the task's session in the bundled runtime, grade the untouched state, close; a verifier that cannot run is named, exit 1 | harbor: one task end to end in Docker, or the exact plan without it; verifiers: import and find the Taskset |
| `benchmarks push <dir>` | `worlds create --from <dir>` (`POST /api/public/worlds`); a new version when the slug exists | `bench push` into the project (the runner's hub) |
| `benchmarks run <slug>` | one hosted run per model: every task × k sessions, graded | one runner job per model: `harbor run` per task, or verifiers `run_eval` in-process |
| `benchmarks results <slug>` | `GET /environments/{id}/performance` | the same endpoint and rows |
| `benchmarks import <dir> --from …` | — | a report of what is expressible as world tasks; writes nothing |

`gateway worlds skill benchmarks-getting-started` prints the same steps in agent form, with expected output at each step. `gateway skill search benchmark` finds it.

## Segmentation

One substrate never runs on the other's infrastructure. Each row is where the rule lives.

| | World benchmark | Container benchmark (harbor, verifiers) |
| - | - | - |
| Project kind | any project; the world page | any project; the environment (benchmark) page |
| Storage | a world version (`POST /api/public/worlds`) | a benchmark-container version (`bench push`: resolve, versions, finalize) |
| Engine | World Host: one hosted session per rollout | benchmark runner (AWS Batch): `harbor run` per task in Docker, or verifiers v1 `run_eval` in-process |
| Results rows | evaluations, samples, `GET /environments/{id}/performance` | the same rows and endpoint |
| Read before writing | `benchmarks push` reads the slug's substrate (`GET /environments/{id}` → `substrate`) and refuses a cross-substrate push with exit 2 and no request; `--kind` that disagrees with the tree is refused before that | the same pre-check; a container benchmark opens on the runner |

A `--kind` that disagrees with the tree is refused by name, with the door that takes the tree. Drop `--kind` to take the tree as it is.

## Scaffold a benchmark

```bash theme={null}
gateway benchmarks init ./crm-bench                                           # world, custom contract
gateway benchmarks init ./slack-bench --connector slack                       # world, a shipped template
gateway benchmarks init ./code-bench --kind harbor --task hello --model openrouter/openai/gpt-5-mini -k 3
gateway benchmarks init ./qa-bench --kind verifiers
```

| Flag | What it does |
| - | - |
| `--kind` | `world` (default), `harbor`, `verifiers`. To benchmark a bare image, package it as a harbor task. |
| `--connector` | World: `custom` (default) or a template from `gateway worlds schema templates`. |
| `--name` | Name and package id. Defaults to the directory name. |
| `--task`, `--prompt` | Id and prompt of the first task (harbor: `instruction.md`). |
| `--image` | Harbor: the task's `docker_image`. Default `python:3.11-slim`. |
| `--model`, `--rollouts`/`-k` | Defaults written to `benchmark.json` for `benchmarks run`. `--model` repeats. |

A harbor tree after `init --kind harbor`:

```
code-bench/
  benchmark.json                     # kind, models, rollouts
  pyproject.toml                     # [project] name + version: the slug and the version push records
  tasks/hello/task.toml              # version, [metadata], [agent] timeout_sec, [verifier] timeout_sec, [environment] docker_image
  tasks/hello/instruction.md         # what the agent is told
  tasks/hello/environment/Dockerfile # the task's container
  tasks/hello/tests/test.sh          # the grader: runs after the agent, writes /logs/verifier/reward.json
  tasks/hello/solution/solve.sh      # the reference solution; must score 1.0
```

The reward contract is the only thing the platform reads: `tests/test.sh` writes `/logs/verifier/reward.json` as named numeric metrics (`reward` is the headline, 1.0 = pass; every other key a sub-reward), or a bare number to `/logs/verifier/reward.txt`. A grader that writes nothing makes the trial a null score, never 0.

## Add tasks

World: a task is a prompt plus verifiers; a verifier is one SQL statement in `verifiers/<name>.sql` returning one number in 0..1 over the world's tables.

```bash theme={null}
gateway worlds schema task add ./crm-bench close_ticket --prompt "Close ticket T-7." --verifier ticket_closed
gateway worlds schema task add ./crm-bench escalate --prompt "..." --verifier state_present --world slack=state_present
```

Harbor: copy `tasks/hello` to `tasks/<name>` and edit `instruction.md`, `tests/test.sh`, `solution/solve.sh` and the image. The task contract in full is on [Benchmarks on worlds](/worlds/benchmarks) and [Harbor benchmarks](/evaluation/benchmarks-harbor).

## Validate

```bash theme={null}
gateway benchmarks validate ./crm-bench
gateway benchmarks validate ./code-bench --json
```

World: the `worlds schema check` report (`tools`, `tasks` with their verifiers, each verifier's `initial_score`, `handler_execution.status`); exit 2 when `gateway-env.toml` holds no `[[tasks]]` entry.

Harbor: per task, `task.toml`, an image (`[environment] docker_image` or `environment/Dockerfile`), a non-empty `instruction.md`, a `tests/test.sh` that writes the reward file, and a note when `solution/solve.sh` is absent; tree-level, `pyproject.toml` and that the runner's own dispatch rule (`detect_env_kind`) reads the tree as harbor. Verifiers: `verifiers>=0.2` pinned, a Taskset exported through `__all__`, reward functions present. Exit 2 on any problem; notes never fail it.

```
./code-bench: harbor benchmark — runner reads it as harbor — ok
  ok   hello: image python:3.11-slim; instruction yes; tests/test.sh yes; solution yes
```

<Info>
  A world verifier whose `initial_score` is `1.0` passes before the agent acts. A harbor `tests/test.sh` that never writes `/logs/verifier/reward.json` grades nothing. Both are named by `validate`.
</Info>

## Dry-run one task

```bash theme={null}
gateway benchmarks dry-run ./crm-bench --task close_ticket
gateway benchmarks dry-run ./code-bench --task hello
gateway benchmarks dry-run ./code-bench --plan --json
gateway benchmarks dry-run ./qa-bench
```

| Flag | What it does |
| - | - |
| `--task` | Which task. Defaults to the first (stderr says so when there are several). |
| `--plan` | Print the plan only; never execute. |
| `--json` | Plan and result as JSON. |

Harbor, with Docker reachable: build `environment/Dockerfile` (or pull `docker_image`), run the container with `tests/` at `/tests`, `solution/` at `/solution` and a scratch `/logs`, execute `bash /solution/solve.sh` then `bash /tests/test.sh`, read `/logs/verifier/reward.json`. That is what the runner does per trial, with the reference solution in place of the agent (harbor's `oracle`). The solution must score 1.0; a lower score or no reward exits 2.

```
ran: docker build -t gateway-benchmarks-dryrun-code-bench-hello -f …/environment/Dockerfile …/environment
ran: docker run --rm -w /app -v …/tests:/tests:ro -v …/solution:/solution:ro -v …/logs:/logs gateway-benchmarks-dryrun-code-bench-hello bash -lc "mkdir -p /logs/verifier && bash /solution/solve.sh && bash /tests/test.sh"
reward: 1 {"reward":1,"file_exists":1,"content_ok":1} · exit 0
```

Without Docker the same plan prints — image, workdir, every mount, both steps, the reward path, the runner's `harbor run` line, the static check of the tests — followed by `not executed: … execution needs Docker — the plan is what would run`, exit 0. `GATEWAY_DOCKER` names the docker binary.

Verifiers: the runner's first pass — import `verifiers.v1`, then the package, find the Taskset in `__all__` — inside Docker: an image built from `python:3.11-slim` + `pip install "verifiers>=0.2"`, then `docker run --rm --network none` with the tree mounted read-only. Importing a package runs its module-level code; by default that happens in the container, never as you. `--here` runs the same import in this machine's Python (`GATEWAY_PYTHON`), as you, with your files and network. Without Docker the plan prints, not executed. The rewards need a model and run in `benchmarks run`. The harbor container also runs with `--network none`, and a task whose `tests/`, `solution/` or `environment/` is a symlink is refused (`docker -v` follows host symlinks).

World: the bundled Python runtime opens the task's session exactly as `worlds serve --task` does (its seed laid, its tools served) with no server left running, grades the untouched state with the task's verifiers, and closes. It prints the four steps, the prompt, tools, verifiers, seed rows per entity, and `reward` / `rewards`; the reward is the untouched world's — an agent's rollout is what `benchmarks run` grades. A verifier that cannot run (bad SQL, a missing entity) is printed as `FAIL verifier '<name>': …` with the database engine's words and exits 1. Without the bundled runtime the steps print and the run is marked not executed, exit 0. A hand-driven rollout: `gateway worlds serve ./crm-bench --task close_ticket --port 8080` then `gateway worlds call http://localhost:8080 <tool> --args '{…}' --grade --close`.

```
plan (world · task close_ticket):
  open      open_world(./crm-bench, task="close_ticket", http=False)
  tools     session.tools
  grade     session.grade()
  close     session.close()
prompt:    Close ticket T-7 and note the reason.
tools:     create_record, get_record, list_records
verifiers: ticket_closed
seed:      records 3
reward: 0 {"ticket_closed":0}
```

## Push

```bash theme={null}
gateway benchmarks push ./crm-bench                    # created crm-bench (mock-tools) <hash> · 2 task(s) · <url>
gateway benchmarks push ./code-bench -m "two tasks"    # pushed code-bench (harbor, benchmark runner) <hash> · 2 task(s) · version <id> · READY
```

| Flag | What it does |
| - | - |
| `--kind` | Refuse unless the tree is this kind. |
| `--slug`, `--name` | Slug and display name. Default to the directory name. |
| `--message`, `-m` | Change reason recorded on the version. |
| `--json` | Print the created world or pushed version as JSON. |

`push` first reads what the slug holds (`GET /api/public/environments/{id}` → `substrate`). One slug holds one substrate: a container tree under a world's slug, or a world tree under a container's slug, is refused with exit 2 and nothing sent (pass `--slug`); the hub's versions door refuses the same push with `benchmark_substrate_mismatch` (409). World, free slug: `POST /api/public/worlds`, the request `gateway worlds create <slug> --from <dir>` sends; the tree is compiled by the schema runtime and refused with the path when the contract does not hold. World, slug that holds a world: the tree lands as the next version through the hub. A directory without `[[tasks]]` is refused. Container: `gateway bench push` into the project — resolve, versions, finalize; a tree that does not validate is refused with the first problems. Later pushes are new content-hashed versions (`gateway bench log`, `diff`, `branch`, `tag`).

## Run: models × tasks × rollouts

```bash theme={null}
gateway benchmarks run code-bench --model openrouter/openai/gpt-5-mini --model anthropic/claude-sonnet-4-6 -k 3
gateway benchmarks run crm-bench -k 5 --no-wait
gateway environment run-status <run_id>
gateway bench run <run_id> --cancel
```

| Flag | What it does |
| - | - |
| `--model`, `-m` | Model, provider-prefixed (`anthropic/…`, `openrouter/vendor/model`, `gateway/vendor/model`); repeats. Defaults to `benchmark.json` in `--dir`. |
| `--rollouts`, `-k` | Rollouts per task, 1–64. Defaults to `benchmark.json`, else 1. |
| `--num-examples`, `-n` | Cap the tasks run. |
| `--version-id` | Pin a `READY` version. Defaults to the newest. |
| `--dir` | Directory whose `benchmark.json` supplies defaults. Defaults to the current directory. |
| `--no-wait`, `--wait-timeout` | Return after dispatch, or the seconds to wait for verdicts (default 1200). |

`run` sends `POST /api/public/environments/{id}/run`, one run per model. A world: each rollout is a session on the World Host graded by the task's verifiers. A harbor tree: the runner runs `harbor run -p tasks -a mini-swe-agent -m <model> -e docker -k <rollouts>` — the agent in each task's container, then `tests/test.sh` — and harvests every trial. A verifiers package: the runner installs it and runs `run_eval` in-process. The verdict prints per model, then per task.

While a run is in flight, `run-status` (and the run page) show its progress from dispatch onward — `0/15 rollouts, 95s since dispatch` while the runner job is queued or starting, harvested counts once it registers its evaluation — and a runner job that died outside the platform settles the run `FAILED` with Batch's reason. `gateway bench run <run_id> --cancel` (or `environment run-status <run_id> --cancel`) sends `POST /api/public/environments/runs/{run_id}/cancel` with the secret key: the run settles `CANCELLED` first, then every runner job it dispatched is terminated, and the answer counts how many AWS Batch accepted (`batch: { jobs, terminated, failed }`). Cancelling twice retries the cleanup (`alreadyCancelled: true`); a run that already settled answers 409 naming its status.

## Read results

```bash theme={null}
gateway benchmarks results code-bench
gateway benchmarks results code-bench --runs --json
```

```
task                          mean reward   pass rate   samples   no score
hello                               1.000        100%         3          0
sort-csv                            0.333         33%         3          0
6 samples, 0 without a score
```

`--runs` lists the runs newest first; `--json` prints the aggregate. Null scores are their own count, never 0. In the app: the world page's or the environment page's **Results** tab; each run under **Evals** links its evaluation and samples.

## Import: what of a container benchmark could be a world task

```bash theme={null}
gateway benchmarks import ./code-bench --from harbor
```

Lists each task with the reason it is or is not expressible as a world task (a harbor task grades a container filesystem; a verifiers task grades completion text; neither is world state) and writes nothing; exit 2 when any task was refused. The tree stays a first-class container benchmark. To grade the same work over world state, author it with `benchmarks init` on a world.

## Where to go next

* [Harbor benchmarks](/evaluation/benchmarks-harbor) — the task layout, the reward contract, the dry-run, the runner.
* [Benchmarks on worlds](/worlds/benchmarks) — the world task contract, verifiers, rollouts.
* [gateway worlds](/cli/worlds) — authoring the world: entities, tools, data, sessions, tests.
* [bench & versions](/cli/bench) — versions, branches, tags, proposals, run configs.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.