Skip to main content
A harbor benchmark grades what an agent leaves behind in a container: files it wrote, a program it fixed, a command it ran. Each task under tasks/ is a container the agent works in, and each ships its own grader. The benchmark runner executes it; the World Host is never involved. The commands are on gateway benchmarks; this page is the contract.

The task layout

gateway benchmarks init ./code-bench --kind harbor writes exactly this, with a hello task whose solution scores 1.0. Add a task by copying the directory.

The reward contract

The platform reads one thing: named numeric metrics in /logs/verifier/reward.json. reward is the headline (1.0 = pass); every other key is a sub-reward shown on the sample. A bare number in /logs/verifier/reward.txt also counts. Any language, any grading method.
A grader that writes no reward file makes the trial a null score, never 0. gateway benchmarks validate names a tests/test.sh that never writes the file; dry-run reports reward: null (ungraded) when it did not.

Validate

Per task: task.toml, an image, a non-empty instruction.md, a tests/test.sh that writes the reward file; solution/solve.sh noted when absent. Tree-level: pyproject.toml, and that the runner’s own dispatch rule reads the tree as harbor — a gateway-env.toml at the root would make it a world to the runner, and validate says so. Exit 2 on any problem.

Dry-run one task

With Docker: build the task’s image, run the container (--network none) with /tests, /solution and a scratch /logs, execute the solution then the grader, read the reward. That is one runner trial with the reference solution in the agent’s place. The printed build and run lines are the exact argv spawned. A task whose tests/, solution/ or environment/ is a symlink is refused: docker -v follows host symlinks. Without Docker: the exact plan — image, workdir, every mount, both steps, the reward path, the runner’s harbor run line, the static check of the tests — and not executed: … execution needs Docker. Exit 2 when the solution scores below 1.0 or writes no reward.

Push, run, results

push is gateway bench push into your project: a content-hashed benchmark-container version (resolve, versions, finalize). run dispatches one runner job per model; the job downloads the version and runs harbor run -p tasks -a mini-swe-agent -m <model> -e docker -k <rollouts> — the agent in each task’s container, then tests/test.sh — and harvests every trial into an evaluation. Models are litellm-style, provider-prefixed; the project’s LLM connection supplies the key. results reads GET /api/public/environments/{id}/performance, the same rows a world benchmark fills.

Segmentation

Where to go next