Skip to main content
A benchmark is one world (tools, data, a schema when the vendor has one) plus tasks plus a run config. Each task is a prompt and the verifiers that grade the world’s end state. A run opens every task k times per model as a session, lets the model act through the world’s tools, grades the end state, and files an evaluation. The result is a per-task aggregate. The commands are on gateway benchmarks. For an agent, gateway worlds skill benchmarks-getting-started prints the same steps.

The task contract

gateway worlds schema task add <dir> <id> --prompt "…" --verifier <name> [--world <alias>=<name>] appends one entry and refuses a verifier file that does not exist, a duplicate id, or an alias that is not linked.

Verifiers

A verifier is one SQL statement over the world’s tables (the entity names in schema/world.json) returning one number in 0..1.
gateway benchmarks check <dir> runs every verifier over the initial state and prints verifiers.<name>.initial_score. A verifier scoring 1.0 before the agent acts grades nothing. The other grader formats live outside the bundle, on a stored task: assertions over entities or a python check. See Keep tasks as project resources.

Dry-run, then grade a task locally before pushing

dry-run opens the task’s session in the bundled runtime the way serve --task does, grades the untouched state with the task’s verifiers, and closes: prompt, tools, verifiers, seed rows, reward and rewards. A verifier that cannot run is named with the database engine’s words and exits 1. serve --task opens one local session on the task for a rollout by hand; call … --grade returns { "reward": 1, "rewards": { "ticket_closed": 1 } } for the end state; --close ends the session.

Run k rollouts per task per model

One hosted run per model. Each run grades every task k times; a rollout is one session, and its rewards carry every verifier’s score. Models are provider-prefixed; the project’s LLM connection supplies the key.
Never set max_tokens on the model. A capped reasoning model returns truncated output and a failed parse, not a low score.

Read the results

Per task: mean reward, pass rate, sample count, and the count of samples with no score. Null scores are their own category and are never coerced to 0. The same table is the world page’s Results tab (…&tab=results); each run under Evals links its evaluation, samples and per-task rewards.

Harbor and verifiers: container benchmarks

A harbor task tree (tasks/<name>/task.toml, instruction.md, tests/test.sh, environment/) grades a container filesystem after the agent runs. A verifiers package (pyproject.toml with verifiers>=0.2, a Taskset exporting reward functions) grades the completion with Python. Both are first-class benchmarks on the benchmark runner, with the same commands. The kind is read from the tree, and push checks it against --kind and against what the slug already holds before writing, so a world and a container benchmark never share a slug. Create and run a container benchmark with gateway benchmarks; it opens on the runner.

Where to go next