tasks/ is a container the agent works in, and each ships its own grader. The benchmark runner executes it; the World Host is never involved. The commands are on gateway benchmarks; this page is the contract.
The task layout
gateway benchmarks init ./code-bench --kind harbor writes exactly this, with a hello task whose solution scores 1.0. Add a task by copying the directory.
The reward contract
The platform reads one thing: named numeric metrics in/logs/verifier/reward.json. reward is the headline (1.0 = pass); every other key is a sub-reward shown on the sample. A bare number in /logs/verifier/reward.txt also counts. Any language, any grading method.
A grader that writes no reward file makes the trial a null score, never 0.
gateway benchmarks validate names a tests/test.sh that never writes the file; dry-run reports reward: null (ungraded) when it did not.Validate
task.toml, an image, a non-empty instruction.md, a tests/test.sh that writes the reward file; solution/solve.sh noted when absent. Tree-level: pyproject.toml, and that the runner’s own dispatch rule reads the tree as harbor — a gateway-env.toml at the root would make it a world to the runner, and validate says so. Exit 2 on any problem.
Dry-run one task
--network none) with /tests, /solution and a scratch /logs, execute the solution then the grader, read the reward. That is one runner trial with the reference solution in the agent’s place. The printed build and run lines are the exact argv spawned. A task whose tests/, solution/ or environment/ is a symlink is refused: docker -v follows host symlinks. Without Docker: the exact plan — image, workdir, every mount, both steps, the reward path, the runner’s harbor run line, the static check of the tests — and not executed: … execution needs Docker. Exit 2 when the solution scores below 1.0 or writes no reward.
Push, run, results
push is gateway bench push into your project: a content-hashed benchmark-container version (resolve, versions, finalize). run dispatches one runner job per model; the job downloads the version and runs harbor run -p tasks -a mini-swe-agent -m <model> -e docker -k <rollouts> — the agent in each task’s container, then tests/test.sh — and harvests every trial into an evaluation. Models are litellm-style, provider-prefixed; the project’s LLM connection supplies the key. results reads GET /api/public/environments/{id}/performance, the same rows a world benchmark fills.
Segmentation
Where to go next
- gateway benchmarks — every command and flag.
- Benchmarks on worlds — the other substrate.
- bench & versions — versions, branches, tags, proposals.