[[tasks]] (a prompt and the SQL verifiers that grade the world’s end state; run on the World Host) or a container benchmark (a harbor task tree graded by its own tests/test.sh inside a container, or a verifiers package graded by Python reward functions; run by the benchmark runner). gateway benchmarks is one set of commands for both. The kind is read from the directory, never guessed.
gateway worlds skill benchmarks-getting-started prints the same steps in agent form, with expected output at each step. gateway skill search benchmark finds it.
Segmentation
One substrate never runs on the other’s infrastructure. Each row is where the rule lives.
A
--kind that disagrees with the tree is refused by name, with the door that takes the tree. Drop --kind to take the tree as it is.
Scaffold a benchmark
A harbor tree after
init --kind harbor:
tests/test.sh writes /logs/verifier/reward.json as named numeric metrics (reward is the headline, 1.0 = pass; every other key a sub-reward), or a bare number to /logs/verifier/reward.txt. A grader that writes nothing makes the trial a null score, never 0.
Add tasks
World: a task is a prompt plus verifiers; a verifier is one SQL statement inverifiers/<name>.sql returning one number in 0..1 over the world’s tables.
tasks/hello to tasks/<name> and edit instruction.md, tests/test.sh, solution/solve.sh and the image. The task contract in full is on Benchmarks on worlds and Harbor benchmarks.
Validate
worlds schema check report (tools, tasks with their verifiers, each verifier’s initial_score, handler_execution.status); exit 2 when gateway-env.toml holds no [[tasks]] entry.
Harbor: per task, task.toml, an image ([environment] docker_image or environment/Dockerfile), a non-empty instruction.md, a tests/test.sh that writes the reward file, and a note when solution/solve.sh is absent; tree-level, pyproject.toml and that the runner’s own dispatch rule (detect_env_kind) reads the tree as harbor. Verifiers: verifiers>=0.2 pinned, a Taskset exported through __all__, reward functions present. Exit 2 on any problem; notes never fail it.
A world verifier whose
initial_score is 1.0 passes before the agent acts. A harbor tests/test.sh that never writes /logs/verifier/reward.json grades nothing. Both are named by validate.Dry-run one task
Harbor, with Docker reachable: build
environment/Dockerfile (or pull docker_image), run the container with tests/ at /tests, solution/ at /solution and a scratch /logs, execute bash /solution/solve.sh then bash /tests/test.sh, read /logs/verifier/reward.json. That is what the runner does per trial, with the reference solution in place of the agent (harbor’s oracle). The solution must score 1.0; a lower score or no reward exits 2.
harbor run line, the static check of the tests — followed by not executed: … execution needs Docker — the plan is what would run, exit 0. GATEWAY_DOCKER names the docker binary.
Verifiers: the runner’s first pass — import verifiers.v1, then the package, find the Taskset in __all__ — inside Docker: an image built from python:3.11-slim + pip install "verifiers>=0.2", then docker run --rm --network none with the tree mounted read-only. Importing a package runs its module-level code; by default that happens in the container, never as you. --here runs the same import in this machine’s Python (GATEWAY_PYTHON), as you, with your files and network. Without Docker the plan prints, not executed. The rewards need a model and run in benchmarks run. The harbor container also runs with --network none, and a task whose tests/, solution/ or environment/ is a symlink is refused (docker -v follows host symlinks).
World: the bundled Python runtime opens the task’s session exactly as worlds serve --task does (its seed laid, its tools served) with no server left running, grades the untouched state with the task’s verifiers, and closes. It prints the four steps, the prompt, tools, verifiers, seed rows per entity, and reward / rewards; the reward is the untouched world’s — an agent’s rollout is what benchmarks run grades. A verifier that cannot run (bad SQL, a missing entity) is printed as FAIL verifier '<name>': … with the database engine’s words and exits 1. Without the bundled runtime the steps print and the run is marked not executed, exit 0. A hand-driven rollout: gateway worlds serve ./crm-bench --task close_ticket --port 8080 then gateway worlds call http://localhost:8080 <tool> --args '{…}' --grade --close.
Push
push first reads what the slug holds (GET /api/public/environments/{id} → substrate). One slug holds one substrate: a container tree under a world’s slug, or a world tree under a container’s slug, is refused with exit 2 and nothing sent (pass --slug); the hub’s versions door refuses the same push with benchmark_substrate_mismatch (409). World, free slug: POST /api/public/worlds, the request gateway worlds create <slug> --from <dir> sends; the tree is compiled by the schema runtime and refused with the path when the contract does not hold. World, slug that holds a world: the tree lands as the next version through the hub. A directory without [[tasks]] is refused. Container: gateway bench push into the project — resolve, versions, finalize; a tree that does not validate is refused with the first problems. Later pushes are new content-hashed versions (gateway bench log, diff, branch, tag).
Run: models × tasks × rollouts
run sends POST /api/public/environments/{id}/run, one run per model. A world: each rollout is a session on the World Host graded by the task’s verifiers. A harbor tree: the runner runs harbor run -p tasks -a mini-swe-agent -m <model> -e docker -k <rollouts> — the agent in each task’s container, then tests/test.sh — and harvests every trial. A verifiers package: the runner installs it and runs run_eval in-process. The verdict prints per model, then per task.
While a run is in flight, run-status (and the run page) show its progress from dispatch onward — 0/15 rollouts, 95s since dispatch while the runner job is queued or starting, harvested counts once it registers its evaluation — and a runner job that died outside the platform settles the run FAILED with Batch’s reason. gateway bench run <run_id> --cancel (or environment run-status <run_id> --cancel) sends POST /api/public/environments/runs/{run_id}/cancel with the secret key: the run settles CANCELLED first, then every runner job it dispatched is terminated, and the answer counts how many AWS Batch accepted (batch: { jobs, terminated, failed }). Cancelling twice retries the cleanup (alreadyCancelled: true); a run that already settled answers 409 naming its status.
Read results
--runs lists the runs newest first; --json prints the aggregate. Null scores are their own count, never 0. In the app: the world page’s or the environment page’s Results tab; each run under Evals links its evaluation and samples.
Import: what of a container benchmark could be a world task
benchmarks init on a world.
Where to go next
- Harbor benchmarks — the task layout, the reward contract, the dry-run, the runner.
- Benchmarks on worlds — the world task contract, verifiers, rollouts.
- gateway worlds — authoring the world: entities, tools, data, sessions, tests.
- bench & versions — versions, branches, tags, proposals, run configs.