> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Harbor benchmarks

> A container benchmark is a harbor task tree — each task a container the agent works in, graded by its own tests. Build one, validate it, dry-run a task, push it to the benchmark runner, run k rollouts per model, read the results.

A harbor benchmark grades what an agent leaves behind in a container: files it wrote, a program it fixed, a command it ran. Each task under `tasks/` is a container the agent works in, and each ships its own grader. The benchmark runner executes it; the World Host is never involved. The commands are on [gateway benchmarks](/cli/benchmarks); this page is the contract.

| Task acts on | Kind | Page |
| - | - | - |
| A vendor's tools, API or data | world | [Benchmarks on worlds](/worlds/benchmarks) |
| A filesystem or a program in a container | harbor | this page |
| The model's completion text, scored in Python | verifiers | [gateway benchmarks](/cli/benchmarks#dry-run-one-task) |

## The task layout

```
code-bench/
  benchmark.json                     # kind, models, rollouts: the defaults `benchmarks run` reads
  pyproject.toml                     # [project] name + version: the slug and the version push records; not a Python package
  tasks/hello/task.toml
  tasks/hello/instruction.md         # what the agent is told
  tasks/hello/environment/Dockerfile # the task's container
  tasks/hello/tests/test.sh          # the grader
  tasks/hello/solution/solve.sh      # the reference solution
```

```toml theme={null}
# tasks/hello/task.toml
version = "1.0"

[metadata]
difficulty = "easy"
category = "smoke"
tags = ["gateway"]

[agent]
timeout_sec = 300.0

[verifier]
timeout_sec = 120.0

[environment]
docker_image = "python:3.11-slim"
```

| Piece | Meaning |
| - | - |
| `instruction.md` | What the agent is told. |
| `[environment] docker_image` or `environment/Dockerfile` | The container the agent works in. A compose file works too. |
| `tests/test.sh` | Runs after the agent, in the same container, with the task's `tests/` at `/tests`. Writes the reward file. |
| `solution/solve.sh` | The reference solution. `benchmarks dry-run` and harbor's `oracle` agent run it; it must score 1.0. The runner ignores it. |
| `[agent] timeout_sec`, `[verifier] timeout_sec` | Wall clocks for the agent and the grader. |

`gateway benchmarks init ./code-bench --kind harbor` writes exactly this, with a `hello` task whose solution scores 1.0. Add a task by copying the directory.

## The reward contract

The platform reads one thing: named numeric metrics in `/logs/verifier/reward.json`. `reward` is the headline (1.0 = pass); every other key is a sub-reward shown on the sample. A bare number in `/logs/verifier/reward.txt` also counts. Any language, any grading method.

```bash theme={null}
#!/bin/bash
mkdir -p /logs/verifier
content_ok=0.0
if [ -f /app/hello.txt ] && [ "$(tr -d '\n' < /app/hello.txt)" = "Hello, Gateway!" ]; then content_ok=1.0; fi
cat > /logs/verifier/reward.json <<EOF
{"reward": ${content_ok}, "content_ok": ${content_ok}}
EOF
```

<Info>
  A grader that writes no reward file makes the trial a null score, never 0. `gateway benchmarks validate` names a `tests/test.sh` that never writes the file; `dry-run` reports `reward: null (ungraded)` when it did not.
</Info>

## Validate

```bash theme={null}
gateway benchmarks validate ./code-bench
```

Per task: `task.toml`, an image, a non-empty `instruction.md`, a `tests/test.sh` that writes the reward file; `solution/solve.sh` noted when absent. Tree-level: `pyproject.toml`, and that the runner's own dispatch rule reads the tree as harbor — a `gateway-env.toml` at the root would make it a world to the runner, and `validate` says so. Exit 2 on any problem.

## Dry-run one task

```bash theme={null}
gateway benchmarks dry-run ./code-bench --task hello
```

With Docker: build the task's image, run the container (`--network none`) with `/tests`, `/solution` and a scratch `/logs`, execute the solution then the grader, read the reward. That is one runner trial with the reference solution in the agent's place. The printed `build` and `run` lines are the exact argv spawned. A task whose `tests/`, `solution/` or `environment/` is a symlink is refused: `docker -v` follows host symlinks. Without Docker: the exact plan — image, workdir, every mount, both steps, the reward path, the runner's `harbor run` line, the static check of the tests — and `not executed: … execution needs Docker`. Exit 2 when the solution scores below 1.0 or writes no reward.

## Push, run, results

```bash theme={null}
gateway benchmarks push ./code-bench -m "two tasks"
gateway benchmarks run code-bench --model openrouter/openai/gpt-5-mini -k 3
gateway benchmarks results code-bench
```

`push` is `gateway bench push` into your project: a content-hashed benchmark-container version (resolve, versions, finalize). `run` dispatches one runner job per model; the job downloads the version and runs `harbor run -p tasks -a mini-swe-agent -m <model> -e docker -k <rollouts>` — the agent in each task's container, then `tests/test.sh` — and harvests every trial into an evaluation. Models are litellm-style, provider-prefixed; the project's LLM connection supplies the key. `results` reads `GET /api/public/environments/{id}/performance`, the same rows a world benchmark fills.

| Thing | Where in the UI |
| - | - |
| The benchmark: Overview, Tasks, Results, Evals, History | `/project/{projectId}/environment?env={containerId}&tab=<tab>` |
| Each run: evaluation, samples, per-task rewards, agent transcript tails | `…&tab=evals` |

## Segmentation

| | World benchmark | Harbor benchmark |
| - | - | - |
| Storage | a world version (`POST /api/public/worlds`) | a benchmark-container version (`bench push`) |
| Engine | World Host: a hosted session per rollout | benchmark runner: `harbor run` per task in Docker |
| Results rows | evaluations, samples, `/environments/{id}/performance` | the same |

## Where to go next

* [gateway benchmarks](/cli/benchmarks) — every command and flag.
* [Benchmarks on worlds](/worlds/benchmarks) — the other substrate.
* [bench & versions](/cli/bench) — versions, branches, tags, proposals.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.