Skip to main content
Every route below lives under /api/public on your Surface Area host and authenticates with your project key pair over HTTP Basic auth. The examples assume GATEWAY_HOST, GATEWAY_PUBLIC_KEY, and GATEWAY_SECRET_KEY are exported, as shown on the Overview page. Paths in the tables omit the /api/public prefix for readability. A path of /traces is $GATEWAY_HOST/api/public/traces.

Worlds

A world is created from a connector template, a published workspace connector, or a tree of files sent as the body. The same request backs gateway worlds create and the MCP tool create_world.

How a world is stored

The platform stores the gateway-world/1 tree as it is, and one session serves both the world’s tools and its HTTP routes. A world’s descriptor carries two scriptable blocks. redact is the connector’s [redact] table — slug (the digest salt, which is the connector’s own slug and can differ from the platform slug), hash, drop, preserve and round — or null when the world declares none. dependencies is the head version’s resolved [dependencies] links, each with its alias, the version it depends on, its task and tools, and the entities mappings that carry a field across. See Link worlds together for what a link does and Keep real data out for what each redaction mode means.

Reach the schema runtime without Python

The world compiler, the store and describe are Python and run on the platform. POST /worlds/runtime reaches them over HTTP: the body is the runtime request verbatim — { "command": "...", "files": [{ "path": "...", "content": "<base64>" }] } — and the answer is the runtime’s own reply, {"status": "ok", ...} or a 400 carrying {"status": "rejected", "error": {code, message, path}} that names the file and field it refused. Allowed commands include validate, compile, init-template, init-template-files, template-summary, describe, import-data, ingest, conform, counts and templates.
The gateway CLI runs the same code on your machine when it finds python3 3.12 or newer, and posts it here when it does not. A gateway worlds schema check on a machine without Python is this route.
List worlds with GET /environments. The world detail and version routes under /environments/{containerId} are documented on the Worlds client page.
The world, world-session and world-task routes read with either key and write with one. A GET accepts Authorization: Bearer $GATEWAY_PUBLIC_KEY as well as Basic auth; every POST, PATCH and DELETE on them requires the secret key over Basic and answers 403 to a public key. The one exception is POST /world-tasks/{taskId}/validate, which writes nothing and takes either.

Put data in a world

Read rows back

Without entity, GET /worlds/{slug}/data answers per-entity row counts. Name an entity and the other parameters open up. Asking for across or compare without an entity answers 400.

Send rows inline

POST /worlds/{slug}/data/import takes { rows, mode?, redact?, message?, dryRun? }, where rows is {"<entity>": [row, ...]}. A large file goes through a batch. dryRun validates every row and reports without writing. The route follows the batch for 55 seconds. It answers 201 with the version the rows became, or 202 with {batchId, state, rowCount} when the batch outlasted the wait — poll GET /worlds/{slug}/data/batches/{batchId} from there. redact chooses what happens to plaintext on the way in: off (the default), apply, or refuse. See Keep real data out for what each mode does.

Upload a large file as a batch

POST /worlds/{slug}/data/batches takes { mode?, merge?, redact?, atomic?, dryRun?, publish?, message?, expectedChunks?, entities? } and answers 201 with { batchId, state: "open", limits }. The limits block holds the sizes to chunk against: chunkBytes (per chunk, compressed), batchBytes (per batch, compressed) and inlineBytes (what the import route takes inline). redact takes the same three values as on import. The sequence is five calls:
  1. Open the batch — POST /worlds/{slug}/data/batches. Keep the batchId and the limits.
  2. Register each chunk — POST …/batches/{batchId}/chunks with { ordinal, digest, bytes, rows, entity? }, where digest is the hex sha256 of the gzipped bytes and bytes is their gzipped size. The answer carries uploadUrl, the uploadHeaders the URL was signed for, and expiresAt an hour out.
  3. Upload the bytes — PUT the gzipped chunk to uploadUrl, sending uploadHeaders exactly as given.
  4. Confirm the upload — POST …/chunks/{chunkId}/uploaded. The server stats the stored object and refuses when its size differs from what was registered.
  5. Seal the batch — POST …/complete answers 202 and state queued. Poll GET …/batches/{batchId} until the state settles.
Registering the same ordinal with the same digest again returns the same chunk, so an interrupted upload resumes. A different digest replaces the chunk while it is unuploaded and is refused once it is not. Bytes the world already holds under that digest, in any of its batches, come back uploaded: true with no URL, so nothing is uploaded twice. complete refuses while any registered chunk is unconfirmed, naming the ordinals, and refuses when the chunk count differs from an expectedChunks the batch declared. Completing a batch that is already queued changes nothing and answers the same body. Confirming a chunk twice answers the same record.

What a batch reports

GET /worlds/{slug}/data/batches/{batchId} answers the batch as the import job sees it: its state, a progress block of {phase, validatedChunks, appliedRows}, entityCounts, the snapshotId and versionId it produced, and an error when it failed. applied, refused and published are final: the job never touches the batch again. A refusal record is {chunk, index, entity, primaryKey, field, pointer, detail} — the chunk’s ordinal, the 0-based line inside it, and the JSON pointer at the value that broke the contract. The batch view carries the first 200 in refusals; reportUrl is a presigned link to every record as JSONL.

Run the tests a world ships

Send {"wait": false} on the POST and the call answers at once with the queued run; read GET /worlds/{slug}/tests/{runId} until it settles. The default holds the request open until the suite finishes.

World sessions

A session is one live copy of a world, pinned to a version and opened on a task. Opening returns immediately with state: PROVISIONING; poll until ready is true, then drive it through /calls. task takes one of three shapes: a bundled task’s name, a stored task reference { id, version? }, or the session’s own spec { instruction, seed?, grader?, metadata?, name? }. Every response carries a task block naming which it was, its content hash, and any stored id and version. surfaces asks for more than tools: api for the world’s service, ui for its dashboard, browser for a hosted browser driving that dashboard. Asking for one the world does not have answers 400 rather than serving fewer. A schema world whose tree carries connector.toml answers surfaces.api from the World Host even when surfaces did not ask for it, because one session serves both surfaces of the same world.
A World Host session’s api surface also carries a token. Send it as Authorization: Bearer <token> in a request header on every call to the API URL — never in the query string.

Stored world tasks

A stored task keeps an instruction, a seed and a grader as a project resource rather than inside a bundle, so many sessions open on it by id and every grade attributes to the task and its version.

Worlds as versioned bundles

Pushing, pinning, branching and dispatching hosted runs live under /benchmark-containers, the older spelling of the same object. See Worlds hub for the Python client and bench & versions for the CLI.

Project files

The project’s file drive, addressed by path. Bytes move over presigned URLs; the app never proxies them. The same routes back the Files tab, gateway files … and the MCP tools list_files, get_file, request_file_upload, confirm_file_upload, delete_file, move_file, create_folder, list_file_connectors and sync_file_connector.
The same drive is reachable per environment at /environments/{containerId}/files… with the same shapes. Writes need a secret key.

Dashboards

A design dashboard’s code: TypeScript render files, an entrypoint and an optional transform. The same routes back gateway dashboards … (Dashboards from the CLI) and the MCP tools list_dashboards, create_dashboard, pull_dashboard, push_dashboard, test_dashboard_transform and preview_dashboard. Every write and every route that runs code needs a secret key.

Traces

Traces are the top-level record of one agent run. List them with filters, or fetch one by id to get its scores and full observation tree.
The list endpoint also takes a fields parameter to control how much of each trace comes back. Request core, io, scores, observations, or metrics as a comma-separated list to trim the payload.

Sessions

A session groups the traces of a multi-turn conversation. List sessions in a time range, or fetch one to walk its traces. With includeIO=true, each trace arrives with its observations, inputs and outputs intact, ready to dump as a test fixture.

Observations

Observations are the steps inside a trace: generations, spans, tool calls, and events. List them with filters or fetch one by id.

Scores

Scores attach an evaluation to a target. Create them one at a time, or list and delete them.
A score targets exactly one thing. Set exactly one of traceId, sessionId, or datasetRunId on the body: never more than one. An observationId may accompany a traceId to point the score at a specific step within that trace.
The dataType selects the value shape: NUMERIC takes a number, CATEGORICAL a string, and BOOLEAN a 0 or 1. Omit dataType to let the value type decide.
The v1 GET /scores list returns trace scores, each carrying its trace’s userId, tags, and environment. Use GET /v2/scores for session and dataset-run scores in the same list.

Datasets, dataset items, and dataset run items

Datasets hold evaluation inputs; runs record what an agent produced against them. The current dataset routes are versioned v2; the unversioned routes remain for backward compatibility.

Annotation queues

Annotation queues route traces and sessions to human reviewers. List queues, then read or fill their items.

Context Hub

Context Hub stores versioned repositories of agent definitions, prompts, skills, and memory. Each kind (agents, prompts, skills, or memory) shares one route shape.
Context Hub items are versioned agent, prompt, skill and memory repositories; traces stamp the items an execution used (see Tracing reference).

User profiles

User profiles store the identity and attributes behind a trace’s userId, tying usage and cost back to a real account. By default a POST merges new attributes into the existing set, so incremental updates are one call. Set mergeAttributes to false to replace the attribute set instead.

User events

User events record the value a user (or the agent acting for them) produced: the per-user ROI stream. value is the event’s worth in your own terms (minutes saved, revenue); Surface Area sums it in the ROI and hours-saved rollups, and evals and task-set replays can assert against the stream.

Agent builds

Agent builds are immutable, content-addressed agent identities: the SDK registers one automatically at tracing.init({ agent, version }), and a release is a label move, never a mutation. Rolling back across a breaking tool change is refused unless forced.

The release gate

The gate is the rule the Releases page applies before a promotion: the pass rate over graded rollouts must reach the project’s threshold, and the graded rollout count must reach its floor (default 0.70 and 30; edit both under Releases). The gate endpoint applies that rule to every completed run attributed to the build, and to other builds registered from the same commit, so a pipeline and the Releases page agree. Missing credentials exit 1, never 2, so a pipeline that waits on untested fails when its secret is not set.
passRate is null (never 0) when nothing was graded. comparedTo is the newest production release before this build, or null when there is none. The CLI wrapper is gateway release gate, whose --wait polls until every run attributed to the build is terminal and whose exit code CI branches on. See The gateway CLI.

Media

Media endpoints handle images, audio, and other binary assets attached to traces and observations. The upload flow requests a URL, uploads to it, then confirms.

Metrics

Metrics endpoints aggregate traces, observations, and scores into numbers. One endpoint runs a flexible query; the other returns daily usage and cost buckets. The query object on GET /metrics takes a view, an array of metrics, optional dimensions, filters, and timeDimension, plus fromTimestamp and toTimestamp. The response is a { "data": [...] } array without a pagination meta block. The metric and dimension element shapes are documented in the Evaluation guide and the Fern API specs.

Models

Model definitions map a model name to its pricing so Surface Area can compute cost. The list mixes your project’s custom models with the built-in global definitions.

More route groups

Surface Area exposes further public route groups: benchmarks (/benchmark-runs, /benchmark-tasks, /benchmark-containers), environments and the environment agent, the coding-agent runtime (/agents, /v1/agents, /agent-runtime), automations, experiment runs (/runs, /experiments), prompt management (/prompts, /v2/prompts), personas, secrets, comments, LLM connections, MCP gateways, integrations (Merge, blob storage, tool sessions, access requests, approvals, webhooks), Slack, and SCIM user and group provisioning. A few administrative route groups (organizations, at /organizations/*, and cross-project management, at /projects) are scoped to an organization key rather than a project key.