> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Worlds hub

> Push, version, browse, and evaluate worlds as content-hashed bundles on Surface Area with the gatewaysdk benchmark hub client.

The hub is where a [world](/worlds) lives as a versioned bundle: push it, pin a run to an exact version, and browse its file tree in the app. The same surface is `gateway bench` on the command line — see [bench & versions](/cli/bench).

<Info>
  The product says **world**; the Python module, the CLI group and the REST
  paths still say `benchmark container`. The words on screen moved to the six in
  the [Glossary](/glossary) while identifiers stayed put, so existing pipelines
  keep working. Read `container` as `world`.
</Info>

A **benchmark container** is a versioned, browsable benchmark environment. The
**source bundle is the versioned source of truth**: every push is content-hashed,
its files are browsable in the app, and every evaluation is pinned to the exact
version it ran against, so scores survive a re-push.

The hub supports **two environment kinds**, auto-detected from the pushed source:

* **verifiers**: a
  [`verifiers`](https://github.com/PrimeIntellect-ai/verifiers)-format Python
  package (`pyproject.toml` + source + a `load_environment()`). The rubric scores
  in-process against any OpenAI-compatible endpoint.
* **harbor**: a directory of Harbor task directories, each grading itself with
  its own script in a sandbox. See [Harbor environments](#harbor-environments)
  below.

The kind is derived structurally: a bundle containing `tasks/*/task.toml` (or
`*/task.toml` at the root) is a harbor environment; everything else is a
verifiers environment.

<Info>
  All commands talk only to your Surface Area host. They read `GATEWAY_HOST`,
  `GATEWAY_PUBLIC_KEY`, and `GATEWAY_SECRET_KEY` from the environment. Writes
  (pushing versions, creating evaluations) require the **secret key**.
</Info>

## Push a container

```bash theme={null}
gateway bench push ./automationbench --visibility PUBLIC
# or: python -m gatewaysdk.benchmark_hub.client push ./automationbench
```

```python theme={null}
from gatewaysdk.benchmark_hub import GatewayBenchmarkClient

client = GatewayBenchmarkClient.from_env()
result = client.push("./automationbench", visibility="PUBLIC")
print(result)
# {containerId, versionId, contentHash, state, slug, alreadyUploaded}
```

The push computes a **source-only content hash** (byte-compatible with Prime
Intellect's hub), uploads a deterministic source tarball, and records the full
file tree for the browser. Re-pushing identical source is a no-op: the existing
version is returned.

## Run a version-pinned evaluation

Harvest an evaluation against a specific container version. Each sample carries
its overall score **and** the named sub-reward breakdown from the rubric.

```python theme={null}
from gatewaysdk.benchmark_hub import GatewayEvalsClient

evals = GatewayEvalsClient.from_env()
eval_id = evals.create_evaluation(version_id, model="gpt-4o-mini")
evals.push_samples(eval_id, samples)            # batched; multi-reward per sample
evals.finalize_evaluation(
    eval_id,
    metrics={"pass_rate": 0.87},
    cost={"total_usd": 0.42},
)
```

## Score with verifiers (one call)

For a `verifiers`-format container, `verifiers` runs the rollout loop against
your own inference, scores it with the rubric, and harvests the result in one
step.

```python theme={null}
import verifiers as vf
from openai import OpenAI
from gatewaysdk.benchmark_hub import GatewayEvalsClient, run_and_harvest

env = vf.load_environment("automationbench")
client = OpenAI(base_url="https://openrouter.ai/api/v1", api_key="...")

run_and_harvest(
    env,
    version_id,
    GatewayEvalsClient.from_env(),
    client=client,
    model="openai/gpt-4o-mini",
    num_examples=100,
    rollouts_per_example=4,        # pass@k + variance
)
```

<Info>
  `verifiers` runs **in-process** against any OpenAI-compatible endpoint
  (including your own), so verifiers containers need no image build.
</Info>

## Harbor environments

A **harbor** environment is a directory of Harbor task directories. Each task
brings its own grader, so any language and any grading method works: pytest, a
SQL check, an LLM judge. The platform never interprets grading logic; it only
reads the numbers the grader writes.

Each task directory holds three required files plus an `environment/` directory:

```
tasks/hello-world/
├── task.toml         # timeouts and the sandbox environment settings
├── instruction.md    # the task given to the agent
├── tests/test.sh     # the grader → writes /logs/verifier/reward.json
└── environment/      # the sandbox definition (e.g. a Dockerfile)
```

Put task directories under a top-level `tasks/` directory, or at the root of the
bundle. A root `pyproject.toml` is optional: when present it supplies only the
name and version for the hub's version tree and enables `--auto-bump`. A
`pyproject.toml` never turns a harbor environment into a verifiers environment.

### The grading contract

The task's `tests/test.sh` runs in the sandbox **after** the agent finishes. It
writes `/logs/verifier/reward.json`, a JSON object of numeric metrics.

```json theme={null}
{ "reward": 1.0, "file_exists": 1.0, "content_ok": 1.0 }
```

The `reward` key is the scalar score for the sample. Every other numeric key
becomes a **named sub-reward**, visible in the evaluation sample drill-down and
averaged across the run. A bare float in `reward.txt` is also accepted as the
scalar score.

<Info>
  The grader writes the reward file: exit codes alone do not grade. A task
  whose `test.sh` exits successfully but writes no reward file scores 0.0.
</Info>

<Info>
  Every harbor task must contain an `environment/` directory (for example
  `environment/Dockerfile`). A `docker_image` in `task.toml` alone is **not**
  sufficient: a task without an `environment/` directory is skipped.
</Info>

### Push a harbor environment

Push is identical to a verifiers container. Point `gateway bench push` at the
directory of task directories.

```bash theme={null}
gateway bench push ./my-env --slug my-env
```

`--auto-bump`, `--rc`, and `--post` work for harbor environments when a root
`pyproject.toml` is present: they rewrite `[project].version` before hashing, so
the bump lands in the bundle and the version tree.

### Evaluate a harbor environment

`gateway bench eval` **auto-detects the kind**: it pulls the bundle and inspects
its file tree. Harbor tasks run in a **local Docker sandbox**, so Docker must be
running.

```bash theme={null}
export OPENROUTER_API_KEY="..."   # read from the environment
gateway bench eval my-env --model openai/gpt-4o-mini
```

The eval runs each task's agent in a sandbox, lets the task's grader score it,
then harvests a version-pinned evaluation with the aggregate reward and each
named sub-reward.

| Flag | Applies to | Default | Purpose |
| - | - | - | - |
| `--kind` | both | `auto` | `auto`, `verifiers`, or `harbor`, to override the detected kind. |
| `--agent` / `-a` | harbor only | `mini-swe-agent` | The agent that runs the tasks (e.g. `claude-code`, or an ACP shorthand like `acp:opencode@1.3.9`). |
| `--model` / `-m` | both | required | The model the agent runs with. |
| `--num-examples` / `-n` | both | all | Cap the number of tasks. |
| `--rollouts-per-example` / `-r` | both | `1` | Attempts per task, for pass\@k and variance. |

<Info>
  Model ids for harbor runs are LiteLLM-style. The CLI adds the `openrouter/`
  prefix automatically when the id does not already carry a provider prefix.
</Info>

### Harvest an existing jobs directory

Harvest a Harbor (or Pier) jobs directory that already ran (no re-execution)
into a version-pinned evaluation with `gateway bench harvest-harbor`.

```bash theme={null}
gateway bench harvest-harbor ./harbor-jobs/my-run my-env@1.0.0 \
  --model openai/gpt-4o-mini
```

The command reads each trial's reward metrics, maps them to evaluation samples,
and pins the result to the resolved container version, the harvest-only twin of
`gateway bench eval --kind harbor`.

<Info>
  A minimal, working harbor environment lives at
  `samples/environments/harbor-hello/`: one task that asks the agent to write a
  file, graded by the task's own `test.sh`. It was pushed, evaluated at reward
  `1.0`, and version-bumped.
</Info>

## Browse in the app

Open **Worlds** in the project sidebar to browse them:

* **Versions**: the content-hashed version history.
* **Code**: the file tree; click a file to view its contents.
* **Evaluations**: every evaluation pinned to the selected version, with sample
  counts and status.

The About sidebar shows a **Kind** badge (verifiers or harbor) derived from the
pushed source.

## Configuration

| Env var | Purpose |
| - | - |
| `GATEWAY_HOST` | Your Surface Area base URL (e.g. `https://withgateway.ai`). |
| `GATEWAY_PUBLIC_KEY` / `GATEWAY_SECRET_KEY` | Project API key pair; writes need the secret key. |

## Where to go next

* [bench & versions](/cli/bench) — the same push, pin, branch and dispatch surface from a terminal.
* [Worlds client](/sdk/environments) — sessions, stored tasks, data and results.
* [What a world is](/worlds) — how a version pin makes two runs comparable.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.