tools/list (or ask your agent to list its tools) to see what is available.
In personal access token (PAT) mode every proxied tool takes an extra
project_id argument that selects the project; it is optional when you set GATEWAY_PROJECT_ID. In project-key (BasicAuth) mode the project is fixed by your keys, so you omit project_id. The “key inputs” below describe each tool’s own arguments; project_id is added by the proxy in PAT mode. A misspelled argument is refused with the names the tool takes.Built-in tools
Two tools are defined by the local proxy itself rather than fetched from the backend.Connectors and worlds
Theconnectors feature authors the contract a world answers to and makes a world from it. A connector template is the vendor connection plus the schema; a world is a built, versioned instance of one.
publish_connector refuses to land over a moved head when you pass
parentCommit, which keeps concurrent agents safe. Validate until it passes,
then publish the same files.connector.toml, its HTTP routes.
World data
Theworld-data feature puts rows into a world or one of its live sessions, reads rows back out, and exports a session’s calls.
where is { "<field>": "<value>" }, and a list of values matches any of them. Omit it for every row of the entity. compare is refused without across, and an unknown entity or field is refused by name.export_world_session, become the rows the next version imports.
Stored world tasks
Theworld-tasks feature keeps a task — instruction, seed and grader — as a versioned project resource, so one task serves many sessions and their grades are comparable.
Open a session with
gateway worlds session open, the
Python client, or POST /api/public/world-sessions,
then pass its id to seed_world_session, export_world_session and
save_world_session_task.World sandboxes
Theworld-sandboxes feature edits a stored world on the platform: open a version, write and run there, check the tree against the push gate, and commit a new version. Each tool is the twin of a gateway worlds sandbox command, and no platform credential enters the sandbox.
To find a sandbox you left open, run
gateway worlds sandbox list.
Dashboards
Thedashboards feature reads and writes a design dashboard’s code: TypeScript render files, an entrypoint and an optional transform. Each tool is the twin of a gateway dashboards command; the same checks as the Design Agent run on every push.
Trace exploration
Theagent-tools feature exposes read tools for analyzing traces, sessions, scores, and project statistics. It is always registered.
execute_query is read-only and accepts SELECT statements only. Provide either sql or queries, never both.Context hub
Thecontext-hub feature gives read-only access to your prompts, skills, memories, and agent definitions — the versioned items an execution can be traced back to. It is always registered.
Environments and task sets
Theenvironments feature lets an agent author, run, and analyze benchmark environments end-to-end. It is registered only in environment-kind projects, so these tools appear in tools/list only when your project is an environment.
- An environment is a versioned benchmark container — a bundle of code plus tasks that you run models against and score.
- A task set is a
kind='tasks'dataset whose items are individual tasks. Tasks are versioned and attributed, exactly like dataset items. - A rollout is one model’s attempt at one task. Every rollout is stored as a Surface Area session, so you can open its full trace.
Read
Write
The full off-platform lifecycle is
create_task_set → create_task (repeat) → link_task_set → run_evaluation → poll list_rollouts / get_rollout_samples → update_task / create_environment_version to iterate. Every tool mirrors the Task Sets and Environment pages in the dashboard one to one.Environment Agent
The Environment Agent is a background analysis agent — one per project — that reads per-task performance, checks reward distributions, and can auto-heal tasks by authoring a new environment version. You can chat with it in the Agent Console or let it run on a schedule or after every evaluation.Console parity: everything the in-app agent can do
Twenty further tools ship with the sameenvironments feature, so they appear under the same environment-project gating. They mount the in-app Environment Agent’s own surface: the same factory, the same handlers, the same durable staging.
Every tool in the five tables below also accepts an optional containerId (an id from list_environments). Pass it whenever the project owns more than one environment; omit it when the project owns one.
Staged edits are keyed to your API key and the environment, and they persist in the database. A file staged in one call is still staged on the next call, from any process.
Stage file edits, then land them
Edits go into a workspace first. Nothing reaches a version until you commit or propose, so batch every file of one coherent change before landing it.env_commit refuses to land over a version the workspace never saw, which
keeps concurrent agents safe. Read env_diff, rebase your edits, and commit
again — or pass acknowledgeDivergence: true to fork deliberately.Review the proposals other agents file
Proposals are also reviewed from here: the apply and dismiss verbs are part of the same tool set.env_apply_proposal refuses when the proposal would overwrite concurrent changes and names the conflicting files. Pass force: true to take the proposal’s side.
Check a version actually runs
A bundle can parse and still be unrunnable. Both tools below block and return the result, so call each once rather than polling.env_run_tests dispatches a real hosted run. A null reward means the verifier graded nothing — a failure to fix,
not a pass.Read a version’s files, reports, and evals
Environment evals are defined in environment code. To change one, edit the source with
env_write_file and land it with env_commit or env_propose.
Schedule analysis and runs
Scheduled work executes in the worker and survives the session closing.Every firing of a
create_run_schedule schedule dispatches one run per
model in the chosen run config. Confirm the environment,
cadence, config, models and pin mode with a person before calling it.
runConfigId may be omitted only when the project has exactly one
schedulable config, and versionId is required for PINNED_BUILD and
refused for LATEST_READY.Prompt management
Theprompts feature adds six tools for managing prompt versions.
Features register by project kind, and a deployment can add modules of its
own.
tools/list is the authority on the tools your instance answers to.Next steps
- Setup — install the client and connect it to your MCP client.
- Getting started with worlds — the same workflow end to end, with each step’s MCP twin named.
- The gateway CLI — the terminal twin of most of the tools above.