gateway bench on the command line — see bench & versions.
The product says world; the Python module, the CLI group and the REST
paths still say
benchmark container. The words on screen moved to the six in
the Glossary while identifiers stayed put, so existing pipelines
keep working. Read container as world.- verifiers: a
verifiers-format Python package (pyproject.toml+ source + aload_environment()). The rubric scores in-process against any OpenAI-compatible endpoint. - harbor: a directory of Harbor task directories, each grading itself with its own script in a sandbox. See Harbor environments below.
tasks/*/task.toml (or
*/task.toml at the root) is a harbor environment; everything else is a
verifiers environment.
All commands talk only to your Surface Area host. They read
GATEWAY_HOST,
GATEWAY_PUBLIC_KEY, and GATEWAY_SECRET_KEY from the environment. Writes
(pushing versions, creating evaluations) require the secret key.Push a container
Run a version-pinned evaluation
Harvest an evaluation against a specific container version. Each sample carries its overall score and the named sub-reward breakdown from the rubric.Score with verifiers (one call)
For averifiers-format container, verifiers runs the rollout loop against
your own inference, scores it with the rubric, and harvests the result in one
step.
verifiers runs in-process against any OpenAI-compatible endpoint
(including your own), so verifiers containers need no image build.Harbor environments
A harbor environment is a directory of Harbor task directories. Each task brings its own grader, so any language and any grading method works: pytest, a SQL check, an LLM judge. The platform never interprets grading logic; it only reads the numbers the grader writes. Each task directory holds three required files plus anenvironment/ directory:
tasks/ directory, or at the root of the
bundle. A root pyproject.toml is optional: when present it supplies only the
name and version for the hub’s version tree and enables --auto-bump. A
pyproject.toml never turns a harbor environment into a verifiers environment.
The grading contract
The task’stests/test.sh runs in the sandbox after the agent finishes. It
writes /logs/verifier/reward.json, a JSON object of numeric metrics.
reward key is the scalar score for the sample. Every other numeric key
becomes a named sub-reward, visible in the evaluation sample drill-down and
averaged across the run. A bare float in reward.txt is also accepted as the
scalar score.
The grader writes the reward file: exit codes alone do not grade. A task
whose
test.sh exits successfully but writes no reward file scores 0.0.Every harbor task must contain an
environment/ directory (for example
environment/Dockerfile). A docker_image in task.toml alone is not
sufficient: a task without an environment/ directory is skipped.Push a harbor environment
Push is identical to a verifiers container. Pointgateway bench push at the
directory of task directories.
--auto-bump, --rc, and --post work for harbor environments when a root
pyproject.toml is present: they rewrite [project].version before hashing, so
the bump lands in the bundle and the version tree.
Evaluate a harbor environment
gateway bench eval auto-detects the kind: it pulls the bundle and inspects
its file tree. Harbor tasks run in a local Docker sandbox, so Docker must be
running.
Model ids for harbor runs are LiteLLM-style. The CLI adds the
openrouter/
prefix automatically when the id does not already carry a provider prefix.Harvest an existing jobs directory
Harvest a Harbor (or Pier) jobs directory that already ran (no re-execution) into a version-pinned evaluation withgateway bench harvest-harbor.
gateway bench eval --kind harbor.
A minimal, working harbor environment lives at
samples/environments/harbor-hello/: one task that asks the agent to write a
file, graded by the task’s own test.sh. It was pushed, evaluated at reward
1.0, and version-bumped.Browse in the app
Open Worlds in the project sidebar to browse them:- Versions: the content-hashed version history.
- Code: the file tree; click a file to view its contents.
- Evaluations: every evaluation pinned to the selected version, with sample counts and status.
Configuration
Where to go next
- bench & versions — the same push, pin, branch and dispatch surface from a terminal.
- Worlds client — sessions, stored tasks, data and results.
- What a world is — how a version pin makes two runs comparable.