Skip to main content
The hub is where a world lives as a versioned bundle: push it, pin a run to an exact version, and browse its file tree in the app. The same surface is gateway bench on the command line — see bench & versions.
The product says world; the Python module, the CLI group and the REST paths still say benchmark container. The words on screen moved to the six in the Glossary while identifiers stayed put, so existing pipelines keep working. Read container as world.
A benchmark container is a versioned, browsable benchmark environment. The source bundle is the versioned source of truth: every push is content-hashed, its files are browsable in the app, and every evaluation is pinned to the exact version it ran against, so scores survive a re-push. The hub supports two environment kinds, auto-detected from the pushed source:
  • verifiers: a verifiers-format Python package (pyproject.toml + source + a load_environment()). The rubric scores in-process against any OpenAI-compatible endpoint.
  • harbor: a directory of Harbor task directories, each grading itself with its own script in a sandbox. See Harbor environments below.
The kind is derived structurally: a bundle containing tasks/*/task.toml (or */task.toml at the root) is a harbor environment; everything else is a verifiers environment.
All commands talk only to your Surface Area host. They read GATEWAY_HOST, GATEWAY_PUBLIC_KEY, and GATEWAY_SECRET_KEY from the environment. Writes (pushing versions, creating evaluations) require the secret key.

Push a container

The push computes a source-only content hash (byte-compatible with Prime Intellect’s hub), uploads a deterministic source tarball, and records the full file tree for the browser. Re-pushing identical source is a no-op: the existing version is returned.

Run a version-pinned evaluation

Harvest an evaluation against a specific container version. Each sample carries its overall score and the named sub-reward breakdown from the rubric.

Score with verifiers (one call)

For a verifiers-format container, verifiers runs the rollout loop against your own inference, scores it with the rubric, and harvests the result in one step.
verifiers runs in-process against any OpenAI-compatible endpoint (including your own), so verifiers containers need no image build.

Harbor environments

A harbor environment is a directory of Harbor task directories. Each task brings its own grader, so any language and any grading method works: pytest, a SQL check, an LLM judge. The platform never interprets grading logic; it only reads the numbers the grader writes. Each task directory holds three required files plus an environment/ directory:
Put task directories under a top-level tasks/ directory, or at the root of the bundle. A root pyproject.toml is optional: when present it supplies only the name and version for the hub’s version tree and enables --auto-bump. A pyproject.toml never turns a harbor environment into a verifiers environment.

The grading contract

The task’s tests/test.sh runs in the sandbox after the agent finishes. It writes /logs/verifier/reward.json, a JSON object of numeric metrics.
The reward key is the scalar score for the sample. Every other numeric key becomes a named sub-reward, visible in the evaluation sample drill-down and averaged across the run. A bare float in reward.txt is also accepted as the scalar score.
The grader writes the reward file: exit codes alone do not grade. A task whose test.sh exits successfully but writes no reward file scores 0.0.
Every harbor task must contain an environment/ directory (for example environment/Dockerfile). A docker_image in task.toml alone is not sufficient: a task without an environment/ directory is skipped.

Push a harbor environment

Push is identical to a verifiers container. Point gateway bench push at the directory of task directories.
--auto-bump, --rc, and --post work for harbor environments when a root pyproject.toml is present: they rewrite [project].version before hashing, so the bump lands in the bundle and the version tree.

Evaluate a harbor environment

gateway bench eval auto-detects the kind: it pulls the bundle and inspects its file tree. Harbor tasks run in a local Docker sandbox, so Docker must be running.
The eval runs each task’s agent in a sandbox, lets the task’s grader score it, then harvests a version-pinned evaluation with the aggregate reward and each named sub-reward.
Model ids for harbor runs are LiteLLM-style. The CLI adds the openrouter/ prefix automatically when the id does not already carry a provider prefix.

Harvest an existing jobs directory

Harvest a Harbor (or Pier) jobs directory that already ran (no re-execution) into a version-pinned evaluation with gateway bench harvest-harbor.
The command reads each trial’s reward metrics, maps them to evaluation samples, and pins the result to the resolved container version, the harvest-only twin of gateway bench eval --kind harbor.
A minimal, working harbor environment lives at samples/environments/harbor-hello/: one task that asks the agent to write a file, graded by the task’s own test.sh. It was pushed, evaluated at reward 1.0, and version-bumped.

Browse in the app

Open Worlds in the project sidebar to browse them:
  • Versions: the content-hashed version history.
  • Code: the file tree; click a file to view its contents.
  • Evaluations: every evaluation pinned to the selected version, with sample counts and status.
The About sidebar shows a Kind badge (verifiers or harbor) derived from the pushed source.

Configuration

Where to go next