Skip to main content
Every tool the Surface Area MCP server exposes, grouped by feature. The proxy discovers tools from your instance at startup, so call tools/list (or ask your agent to list its tools) to see what is available.
In personal access token (PAT) mode every proxied tool takes an extra project_id argument that selects the project; it is optional when you set GATEWAY_PROJECT_ID. In project-key (BasicAuth) mode the project is fixed by your keys, so you omit project_id. The “key inputs” below describe each tool’s own arguments; project_id is added by the proxy in PAT mode. A misspelled argument is refused with the names the tool takes.

Built-in tools

Two tools are defined by the local proxy itself rather than fetched from the backend.

Connectors and worlds

The connectors feature authors the contract a world answers to and makes a world from it. A connector template is the vendor connection plus the schema; a world is a built, versioned instance of one.
publish_connector refuses to land over a moved head when you pass parentCommit, which keeps concurrent agents safe. Validate until it passes, then publish the same files.
A world is stored as its schema tree. One session serves its tools and, when the tree has a connector.toml, its HTTP routes.

World data

The world-data feature puts rows into a world or one of its live sessions, reads rows back out, and exports a session’s calls.
where is { "<field>": "<value>" }, and a list of values matches any of them. Omit it for every row of the entity. compare is refused without across, and an unknown entity or field is refused by name.
A real run’s calls, exported with export_world_session, become the rows the next version imports.

Stored world tasks

The world-tasks feature keeps a task — instruction, seed and grader — as a versioned project resource, so one task serves many sessions and their grades are comparable.
Open a session with gateway worlds session open, the Python client, or POST /api/public/world-sessions, then pass its id to seed_world_session, export_world_session and save_world_session_task.

World sandboxes

The world-sandboxes feature edits a stored world on the platform: open a version, write and run there, check the tree against the push gate, and commit a new version. Each tool is the twin of a gateway worlds sandbox command, and no platform credential enters the sandbox. To find a sandbox you left open, run gateway worlds sandbox list.

Dashboards

The dashboards feature reads and writes a design dashboard’s code: TypeScript render files, an entrypoint and an optional transform. Each tool is the twin of a gateway dashboards command; the same checks as the Design Agent run on every push.

Trace exploration

The agent-tools feature exposes read tools for analyzing traces, sessions, scores, and project statistics. It is always registered.
execute_query is read-only and accepts SELECT statements only. Provide either sql or queries, never both.

Context hub

The context-hub feature gives read-only access to your prompts, skills, memories, and agent definitions — the versioned items an execution can be traced back to. It is always registered.

Environments and task sets

The environments feature lets an agent author, run, and analyze benchmark environments end-to-end. It is registered only in environment-kind projects, so these tools appear in tools/list only when your project is an environment.
  • An environment is a versioned benchmark container — a bundle of code plus tasks that you run models against and score.
  • A task set is a kind='tasks' dataset whose items are individual tasks. Tasks are versioned and attributed, exactly like dataset items.
  • A rollout is one model’s attempt at one task. Every rollout is stored as a Surface Area session, so you can open its full trace.

Read

Write

The full off-platform lifecycle is create_task_set → create_task (repeat) → link_task_set → run_evaluation → poll list_rollouts / get_rollout_samples → update_task / create_environment_version to iterate. Every tool mirrors the Task Sets and Environment pages in the dashboard one to one.

Environment Agent

The Environment Agent is a background analysis agent — one per project — that reads per-task performance, checks reward distributions, and can auto-heal tasks by authoring a new environment version. You can chat with it in the Agent Console or let it run on a schedule or after every evaluation.

Console parity: everything the in-app agent can do

Twenty further tools ship with the same environments feature, so they appear under the same environment-project gating. They mount the in-app Environment Agent’s own surface: the same factory, the same handlers, the same durable staging. Every tool in the five tables below also accepts an optional containerId (an id from list_environments). Pass it whenever the project owns more than one environment; omit it when the project owns one. Staged edits are keyed to your API key and the environment, and they persist in the database. A file staged in one call is still staged on the next call, from any process.

Stage file edits, then land them

Edits go into a workspace first. Nothing reaches a version until you commit or propose, so batch every file of one coherent change before landing it.
env_commit refuses to land over a version the workspace never saw, which keeps concurrent agents safe. Read env_diff, rebase your edits, and commit again — or pass acknowledgeDivergence: true to fork deliberately.

Review the proposals other agents file

Proposals are also reviewed from here: the apply and dismiss verbs are part of the same tool set. env_apply_proposal refuses when the proposal would overwrite concurrent changes and names the conflicting files. Pass force: true to take the proposal’s side.

Check a version actually runs

A bundle can parse and still be unrunnable. Both tools below block and return the result, so call each once rather than polling.
env_run_tests dispatches a real hosted run. A null reward means the verifier graded nothing — a failure to fix, not a pass.

Read a version’s files, reports, and evals

Environment evals are defined in environment code. To change one, edit the source with env_write_file and land it with env_commit or env_propose.

Schedule analysis and runs

Scheduled work executes in the worker and survives the session closing.
Every firing of a create_run_schedule schedule dispatches one run per model in the chosen run config. Confirm the environment, cadence, config, models and pin mode with a person before calling it. runConfigId may be omitted only when the project has exactly one schedulable config, and versionId is required for PINNED_BUILD and refused for LATEST_READY.

Prompt management

The prompts feature adds six tools for managing prompt versions.
Features register by project kind, and a deployment can add modules of its own. tools/list is the authority on the tools your instance answers to.

Next steps