Skip to main content
gateway bench treats a world as a versioned bundle. Every push is content-hashed and browsable in the app, every run records the version it ran against, and any version stays addressable afterwards. Building a benchmark? A benchmark is a world with tasks; gateway benchmarks runs the worlds commands in benchmark order (init, check, push, run, results) and uses bench underneath for versions.
The product says world; the command group is still called bench and its arguments still say container. The words on screen moved to the six in the Glossary while identifiers stayed put, so existing pipelines keep working. Read container as world.

Push a version

bench push uploads the directory as a new version on a branch. The bundle is content-hashed, so pushing an unchanged tree lands on the existing version instead of creating a duplicate.
A world is pushed as the tree it is: schema/world.json, connector.toml, tasks and verifiers. The platform compiles and hosts that tree, so nothing is repackaged.
Pushing onto a branch whose tip moved since you pulled is refused with a 409. Pull or checkout the new tip and push again. Use --force only when you mean to discard what moved.

Address a version with slug@ref

Most read commands take slug[@ref]. A ref resolves in four ways. A bare slug means the head. bench pull also accepts @latest, and bench checkout reads a bare slug as main.
A branch moves; a tag does not. Pin a release gate or a published demo to a tag, and let CI track a branch.

Get a version onto your disk

Use checkout when you intend to push back, pull when you only want to read a version’s files.
Pulling is for authoring. To run an agent against a world, open a hosted session by slug instead of unwrapping a copy on your own disk; see gateway worlds.

Read the history

Branch, tag and fork

Fork when a world diverges permanently — a second tenant, a different vendor plan. Branch when the change is meant to come back.

Propose a change instead of pushing one

bench propose files the directory’s local edits as a change proposal against a branch, which someone then reviews and applies. Use it where a push would be a pull request: an agent editing a world it does not own, or a change that needs a human before it reaches main.
propose takes --title, --message/-m (aliased --reason), and --branch/-b for the branch it merges into (default main). proposals <slug> lists them and filters with --status open|applied|dismissed.

Dispatch hosted runs

A run config is a saved recipe — models, tasks, parallelism — that the platform executes on its own runners. Dispatching one from the CLI sends the same request as the Run button.
dispatch returns as soon as the platform accepts the runs. bench run <runId> --wait polls one to terminal. run-configs create takes --name, --models (comma-separated), --parallelism (1–16), --tasks (bundled task names) or --world-tasks (stored task refs as <id> or <id>@<version>, for api worlds and not alongside --tasks), --agent for a harbor agent, and --auto-run-on-push to dispatch on every pushed READY version. delete takes --id. bench validate combines the three CI steps: --run-config <name> names the config to dispatch, --min-score fails any run below the bar, --summary-file writes a markdown summary for a pull-request comment, and --slug, --branch and --timeout-min control the rest. For a world whose agent runs locally, use gateway worlds ci instead.
Secrets are named in UPPER_SNAKE_CASE and scoped to one world with --env. Omit --value on set and the value is read from standard input, which keeps it out of shell history.

Read the tests a world ships

bench tests <slug> lists the runs of a world’s own test suite, newest first — those triggered on push, from the CLI and from the API alike.
--run <id> prints one run’s full report instead of the list. Authoring and running the suite is gateway worlds test.

Commands that need the Python package

bench eval (local verifiers and harbor engines), bench install (pip-installs a version’s wheel), bench harvest-harbor (harvests a local harbor jobs directory) and bench push --build-wheel execute Python locally and belong to the Python package — see the Python SDK overview. For a hosted run, use bench dispatch.

Where to go next