> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Benchmarks on worlds

> A benchmark is a world with tasks. Write a task as a prompt plus SQL verifiers over the world's state, run every task k times per model, and read the per-task results.

A benchmark is one [world](/worlds) (tools, data, a schema when the vendor has one) plus tasks plus a run config. Each task is a prompt and the verifiers that grade the world's end state. A run opens every task k times per model as a session, lets the model act through the world's tools, grades the end state, and files an evaluation. The result is a per-task aggregate.

| Piece | File or command |
| - | - |
| The world | `schema/world.json`, `gateway-env.toml` with `[schema]`, `tools/handlers.py`, optional `connector.toml` |
| Its data | `data/initial.json`; `gateway worlds data import` |
| Its tasks | `[[tasks]]` in `gateway-env.toml`; `gateway worlds schema task add` |
| The graders | `verifiers/<name>.sql`, one number in 0..1 each |
| The run | `gateway benchmarks run <slug> --model … -k <rollouts>` |
| The results | `gateway benchmarks results <slug>`; the world page's **Results** tab |

The commands are on [gateway benchmarks](/cli/benchmarks). For an agent, `gateway worlds skill benchmarks-getting-started` prints the same steps.

## The task contract

```toml theme={null}
# gateway-env.toml
[[tasks]]
id = "close_ticket"
prompt = "Close ticket T-7 and record the reason as 'duplicate'."
verifiers = ["ticket_closed", "reason_recorded"]

[[tasks]]
id = "escalate"
prompt = "Escalate the locked account to #security on Slack."
verifiers = ["state_present"]
worlds = { slack = ["message_posted"] }
```

| Field | Meaning |
| - | - |
| `id` | The task name; `[A-Za-z_][A-Za-z0-9_-]*`. What `worlds serve --task` and the results table use. |
| `prompt` | What the agent is told. |
| `verifiers` | Names of `verifiers/<name>.sql` in this world. Each becomes a key in `rewards`; `reward` is their mean. |
| `worlds` | A linked world's verifiers, graded on that world's own state; the key is `<alias>/<name>`. Needs `[dependencies.<alias>]` (`gateway worlds schema link`). |

`gateway worlds schema task add <dir> <id> --prompt "…" --verifier <name> [--world <alias>=<name>]` appends one entry and refuses a verifier file that does not exist, a duplicate id, or an alias that is not linked.

## Verifiers

A verifier is one SQL statement over the world's tables (the entity names in `schema/world.json`) returning one number in 0..1.

```sql theme={null}
-- verifiers/ticket_closed.sql
SELECT CASE WHEN COUNT(*) > 0 THEN 1.0 ELSE 0.0 END
FROM "crm_tickets" WHERE id = 'T-7' AND status = 'closed';
```

`gateway benchmarks check <dir>` runs every verifier over the initial state and prints `verifiers.<name>.initial_score`. A verifier scoring `1.0` before the agent acts grades nothing.

The other grader formats live outside the bundle, on a stored task: `assertions` over entities or a `python` check. See [Keep tasks as project resources](/cli/worlds#keep-tasks-as-project-resources).

## Dry-run, then grade a task locally before pushing

```bash theme={null}
gateway benchmarks dry-run ./crm-bench --task close_ticket
gateway worlds serve ./crm-bench --task close_ticket --port 8080 &
gateway worlds call http://localhost:8080 update_ticket --args '{"id":"T-7","status":"closed"}' --grade --close
```

`dry-run` opens the task's session in the bundled runtime the way `serve --task` does, grades the untouched state with the task's verifiers, and closes: prompt, tools, verifiers, seed rows, `reward` and `rewards`. A verifier that cannot run is named with the database engine's words and exits 1. `serve --task` opens one local session on the task for a rollout by hand; `call … --grade` returns `{ "reward": 1, "rewards": { "ticket_closed": 1 } }` for the end state; `--close` ends the session.

## Run k rollouts per task per model

```bash theme={null}
gateway benchmarks push ./crm-bench
gateway benchmarks run crm-bench --model openai/gpt-5-mini --model anthropic/claude-sonnet-4-6 -k 3
```

One hosted run per model. Each run grades every task k times; a rollout is one session, and its `rewards` carry every verifier's score. Models are provider-prefixed; the project's LLM connection supplies the key.

<Info>
  Never set `max_tokens` on the model. A capped reasoning model returns truncated output and a failed parse, not a low score.
</Info>

## Read the results

```bash theme={null}
gateway benchmarks results crm-bench
gateway benchmarks results crm-bench --runs --json
```

Per task: mean reward, pass rate, sample count, and the count of samples with no score. Null scores are their own category and are never coerced to 0. The same table is the world page's **Results** tab (`…&tab=results`); each run under **Evals** links its evaluation, samples and per-task rewards.

## Harbor and verifiers: container benchmarks

A harbor task tree (`tasks/<name>/task.toml`, `instruction.md`, `tests/test.sh`, `environment/`) grades a container filesystem after the agent runs. A verifiers package (`pyproject.toml` with `verifiers>=0.2`, a Taskset exporting reward functions) grades the completion with Python. Both are first-class benchmarks on the benchmark runner, with the same commands.

| You have | Do |
| - | - |
| A task that acts through tools or an API | Author it on a world: `gateway benchmarks init`, `gateway worlds schema task add`. |
| A task that acts on a filesystem or a program | `gateway benchmarks init <dir> --kind harbor`, then `validate`, `dry-run`, `push`, `run`, `results`: [Harbor benchmarks](/evaluation/benchmarks-harbor). |
| A harbor tree or a verifiers package you already have | `gateway benchmarks validate <dir>`, `dry-run`, `push`; the runner executes and grades it, and `benchmarks results` reads it the same way. |
| Both, and you want to know what could move | `gateway benchmarks import <dir> --from harbor\|verifiers` lists each task with the reason it is or is not expressible as a world task. It writes nothing. |

The kind is read from the tree, and `push` checks it against `--kind` and against what the slug already holds before writing, so a world and a container benchmark never share a slug. Create and run a container benchmark with `gateway benchmarks`; it opens on the runner.

## Where to go next

* [gateway benchmarks](/cli/benchmarks) — every command and flag.
* [Getting started with worlds](/worlds/getting-started) — the world itself.
* [Link worlds together](/worlds/links) — tasks graded on two worlds.
* [Harbor benchmarks](/evaluation/benchmarks-harbor) — the container substrate.
* [Evaluation](/evaluation) — judges over runs, and the boundary with benchmarks.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.