The commands are on gateway benchmarks. For an agent,
gateway worlds skill benchmarks-getting-started prints the same steps.
The task contract
gateway worlds schema task add <dir> <id> --prompt "…" --verifier <name> [--world <alias>=<name>] appends one entry and refuses a verifier file that does not exist, a duplicate id, or an alias that is not linked.
Verifiers
A verifier is one SQL statement over the world’s tables (the entity names inschema/world.json) returning one number in 0..1.
gateway benchmarks check <dir> runs every verifier over the initial state and prints verifiers.<name>.initial_score. A verifier scoring 1.0 before the agent acts grades nothing.
The other grader formats live outside the bundle, on a stored task: assertions over entities or a python check. See Keep tasks as project resources.
Dry-run, then grade a task locally before pushing
dry-run opens the task’s session in the bundled runtime the way serve --task does, grades the untouched state with the task’s verifiers, and closes: prompt, tools, verifiers, seed rows, reward and rewards. A verifier that cannot run is named with the database engine’s words and exits 1. serve --task opens one local session on the task for a rollout by hand; call … --grade returns { "reward": 1, "rewards": { "ticket_closed": 1 } } for the end state; --close ends the session.
Run k rollouts per task per model
rewards carry every verifier’s score. Models are provider-prefixed; the project’s LLM connection supplies the key.
Never set
max_tokens on the model. A capped reasoning model returns truncated output and a failed parse, not a low score.Read the results
…&tab=results); each run under Evals links its evaluation, samples and per-task rewards.
Harbor and verifiers: container benchmarks
A harbor task tree (tasks/<name>/task.toml, instruction.md, tests/test.sh, environment/) grades a container filesystem after the agent runs. A verifiers package (pyproject.toml with verifiers>=0.2, a Taskset exporting reward functions) grades the completion with Python. Both are first-class benchmarks on the benchmark runner, with the same commands.
The kind is read from the tree, and
push checks it against --kind and against what the slug already holds before writing, so a world and a container benchmark never share a slug. Create and run a container benchmark with gateway benchmarks; it opens on the runner.
Where to go next
- gateway benchmarks — every command and flag.
- Getting started with worlds — the world itself.
- Link worlds together — tasks graded on two worlds.
- Harbor benchmarks — the container substrate.
- Evaluation — judges over runs, and the boundary with benchmarks.