# Surface Area > Worlds, tracing and evaluation for agents: run agents against reality before reality matters. - [Introduction](https://docs.surfacearea.ai/index.md): Surface Area rebuilds the systems your agent works against as a sealed world, so you can run a candidate version against reality before reality matters. - [Glossary](https://docs.surfacearea.ai/glossary.md): The six words Surface Area uses for the things it runs - world, scenario, rollout, run, eval, and gate. - [Platform overview](https://docs.surfacearea.ai/get-started/platform-overview.md): What Surface Area does, the pieces it is made of, the ways you can drive it, and where to start for the job you have. - [Get Started with Surface Area](https://docs.surfacearea.ai/get-started/index.md): One sitting, end to end - install the CLI, sign in, create a world from a connector, open a session, point an agent at it, grade the run, and read the result in the dashboard. - [Installation](https://docs.surfacearea.ai/get-started/installation.md): Install the gateway command-line tool, sign in to a project, and install the Python SDK when you want to trace an agent. - [Start with your coding agent](https://docs.surfacearea.ai/get-started/agent-prompts.md): Install the platform's guides into your repository, then paste a prompt into Claude Code, Cursor or Codex to have it set up worlds, scenarios, tracing and CI for you. - [Quickstart](https://docs.surfacearea.ai/get-started/quickstart.md): Create a world, open a hosted session, call it over HTTP, and grade the result - in under ten minutes with the gateway command-line tool. - [Dashboard Basics](https://docs.surfacearea.ai/get-started/dashboard.md): Sign in, find your organization and project, create API keys, and read the project landing page. - [Use cases](https://docs.surfacearea.ai/use-cases/index.md): Two recipes built on worlds - gate every pull request on a graded run, and hand each of your own users a sealed simulation seeded from their own data. - [Evals in CI/CD](https://docs.surfacearea.ai/use-cases/evals-in-ci.md): A GitHub Actions workflow that opens a session per scenario on the platform, runs the agent under test against each session's URL, grades the end state, and fails the pull request below a threshold. - [Simulations for your users, from their real data](https://docs.surfacearea.ai/use-cases/sims-from-real-data.md): The funnel from real traffic to a graded per-user run - get the rows out, redact them on the way in, give each tenant a world or each user a task, seed the session, run, grade, and look a user up again afterwards. - [Worlds](https://docs.surfacearea.ai/worlds/index.md): A world is a sealed replica of the systems your agent works against, built from real sessions and pinned to a version, so a candidate can be run against it as often as you like. - [Getting started with worlds](https://docs.surfacearea.ai/worlds/getting-started.md): Build your first world end to end. Authenticate, start from a connector or a contract, put data in, publish a version, open a session, run an agent against it, and iterate. - [Mock any vendor API](https://docs.surfacearea.ai/worlds/custom-connections.md): Describe a third-party API your agent calls, capture what it really returns, seed a world from it through the schema, and let the world serve the vendor's routes so the agent runs against it unchanged. - [The files a world is made of](https://docs.surfacearea.ai/worlds/contract.md): A reference for the three files you edit by hand — schema/world.json and its x-gateway block, gateway-env.toml, and connector.toml — with every key, its default, and what the compiler refuses. - [Give a world a UI](https://docs.surfacearea.ai/worlds/ui.md): Every schema world with routes can serve a UI. Declare it in connector.toml, let compile generate it from the contract or bring a frontend built with any framework, then open a session with the ui and browser surfaces. - [Put data in a world](https://docs.surfacearea.ai/worlds/data.md): Load rows into a world from a file, seed a live session, or pull a real run's calls, transform them with your own code or a declared mapping, and ingest them. One gate, the world's contract, on every route. - [Keep real data out of a shared world](https://docs.surfacearea.ai/worlds/redaction.md): Declare a [redact] policy so a world built from real vendor data can be shared. The four modes, the rules that decide which one a field falls under, the salt, and how to redact rows on the way in. - [Link worlds and resolve entities across them](https://docs.surfacearea.ai/worlds/links.md): Declare how one world's entities map onto another's, normalize values that are not stored the same way, and ask one question across every linked world from the plaintext you know. - [Test a world](https://docs.surfacearea.ai/worlds/tests.md): A world ships its own tests. Write HTTP cases, tool cases and pytest files under tests/, run them locally, on the platform, or on every push, and read the report grouped by route. - [Benchmarks on worlds](https://docs.surfacearea.ai/worlds/benchmarks.md): A benchmark is a world with tasks. Write a task as a prompt plus SQL verifiers over the world's state, run every task k times per model, and read the per-task results. - [What makes a world good](https://docs.surfacearea.ai/worlds/good-world.md): Seven properties that separate a world your agent can be trusted against from a mock that flatters it, each with a check you can run, and a definition of done. - [Spin worlds up and down](https://docs.surfacearea.ai/worlds/sessions.md): A session is one running copy of a pinned world version. Open it, point your agent at its URL, grade it, and close it — or pin it so the URL stays up. - [Run many sessions at once](https://docs.surfacearea.ai/worlds/parallel.md): Every session is its own clone of a pinned world version, so N sessions run side by side without leaking state. Open them from a shell loop, from Python, or from a CI job. - [Simulations for your users](https://docs.surfacearea.ai/worlds/for-your-users.md): Serve a sealed replica to each of your own customers — a world per tenant, or one world with a task per end user — provisioned from one API call and graded per run. - [Native computer use (Anthropic, OpenAI)](https://docs.surfacearea.ai/worlds/computer-use.md): Hand a world session's browser to Claude's or OpenAI's own computer-use tool. The model sees screenshots and clicks at pixel coordinates; the SDK runs each action on the session's browser, returns the screen in the shape the provider expects, and traces every step with its screenshot. - [Run API worlds in CI](https://docs.surfacearea.ai/worlds/ci.md): A GitHub Actions workflow that runs each world's own test suite on every pull request and smoke-tests your agent against a live session, with the full YAML to copy. - [Build a Slack world](https://docs.surfacearea.ai/worlds/slack.md): Create a populated Slack world from synthetic data or selected conversations, run an agent against an immutable version, and evaluate a collaboration scenario. - [Author worlds from chat](https://docs.surfacearea.ai/worlds/assistant.md): The in-dashboard Assistant can author a connector, validate it, publish it, build a world from it, import data, and open a session - without you leaving the chat. - [The sample repository](https://docs.surfacearea.ai/worlds/sample-repo.md): connector-worlds is a public demo app - an example agent (Relay) plus three commands that turn a user's production traffic into hosted worlds on the platform and run the agent against them. - [Evaluation (LLM-as-Judge)](https://docs.surfacearea.ai/evaluation/index.md): How a world's runs are scored - write the criteria a judge grades against, point a check at the runs you care about, and read the scores it produces. - [Agent Replay](https://docs.surfacearea.ai/evaluation/agent-replay.md): Turn real production sessions into a regression suite - replay them against a new agent version and read the pass, fail, and drift results as sessions and scored runs in Surface Area. - [Scenarios & task sets](https://docs.surfacearea.ai/evaluation/data-task-sets.md): Collect real traces and sessions into datasets, mark golden regression sets, and curate the scenarios a world runs. - [Harbor benchmarks](https://docs.surfacearea.ai/evaluation/benchmarks-harbor.md): A container benchmark is a harbor task tree — each task a container the agent works in, graded by its own tests. Build one, validate it, dry-run a task, push it to the benchmark runner, run k rollouts per model, read the results. - [What is tracing](https://docs.surfacearea.ai/tracing/index.md): Capture what your agent did (every LLM call, tool call, and step) and send it to Surface Area for inspection and evaluation. - [Instrument an agent](https://docs.surfacearea.ai/tracing/setup.md): Initialize the Surface Area SDK with one call, then capture spans with the decorator and context-manager APIs. - [Sessions & traces](https://docs.surfacearea.ai/tracing/sessions.md): Group related traces into one session so a multi-turn conversation or workflow reads as a single interaction. - [Auto-instrumentation](https://docs.surfacearea.ai/tracing/auto-instrumentation.md): Capture LLM and agent-framework calls automatically: the frameworks Surface Area instruments and how to control which ones are active. - [Metadata & identity](https://docs.surfacearea.ai/tracing/metadata.md): Attach the agent, user, session, tags, and custom attributes that Surface Area displays and filters on. - [Export to Surface Area](https://docs.surfacearea.ai/tracing/export.md): Configure where spans go over OTLP, how they are batched, and how to flush them reliably. - [Bring traces from any source](https://docs.surfacearea.ai/tracing/any-source.md): Send OpenTelemetry traces from Braintrust, Galileo, Arize/Phoenix, LangSmith, or any OTLP exporter to Surface Area, and link them to sessions and worlds. - [Hooks for any agent](https://docs.surfacearea.ai/hooks/index.md): One rules file around an agent's tool calls — deny, rewrite, add context, record, and put a persona in front of every private individual — attached to any agent SDK, or run as a coding agent's hook. - [Hooks in Claude Code and coding agents](https://docs.surfacearea.ai/hooks/coding-agents.md): Run gateway-hooks.toml as a coding agent's hook — one settings entry, the agent's own event and answer contract, personas that survive across the process the agent spawns per event. - [Hooks with any agent SDK](https://docs.surfacearea.ai/hooks/agent-sdks.md): Attach gateway-hooks.toml to the tools of whichever agent SDK you build on — OpenAI Agents SDK, Vercel AI SDK, a hand-rolled Anthropic tool loop, LangChain, an MCP server, the Claude Agent SDK — by wrapping the functions once. - [The gateway CLI](https://docs.surfacearea.ai/cli/index.md): Install the gateway command from the @withgateway/sdk npm package, sign in to a project, and learn the conventions every subcommand shares — credentials, the local schema runtime, output, and exit codes. - [gateway worlds](https://docs.surfacearea.ai/cli/worlds.md): Every gateway worlds subcommand, grouped by the workflow it belongs to — author a contract, put data in, publish a version, open live sessions, run an agent, and test the world in CI. - [gateway benchmarks](https://docs.surfacearea.ai/cli/benchmarks.md): Build, validate, dry-run, push, run and read a benchmark from the terminal. A benchmark is a world with tasks or a container benchmark (harbor, verifiers); one set of commands, the kind read from the tree. - [gateway bench](https://docs.surfacearea.ai/cli/bench.md): Push a world as a content-hashed version, address any version with slug@ref, branch and tag it, file proposals instead of pushes, and dispatch hosted runs from the terminal. - [gateway files](https://docs.surfacearea.ai/cli/files.md): Upload, list, download, move and delete files in the project's drive from the terminal, fetch a URL into it, sync its connectors, and use drive: as an input or output of any worlds data command. - [gateway traces](https://docs.surfacearea.ai/cli/traces.md): List, inspect and export the project's traces, observations and trace sessions from the terminal, land an export in the drive, and import OTLP/JSON exports from other tools. - [Dashboards from the CLI](https://docs.surfacearea.ai/cli/dashboards.md): Create a design dashboard, check its TypeScript render code out into a directory, edit it with any editor or coding agent, push it back and render it on real data with gateway dashboards. - [Python SDK](https://docs.surfacearea.ai/sdk/index.md): What the gatewaysdk Python package gives you — a client for worlds, their sessions and their versions, plus runs and scoring — and which module to import for each task. - [Worlds client](https://docs.surfacearea.ai/sdk/environments.md): Open live world sessions, store tasks, stream rows in, and read a world's scenarios, per-task performance and rollouts — the gatewaysdk clients for worlds and the runs made against them. - [Worlds hub](https://docs.surfacearea.ai/sdk/benchmark-hub.md): Push, version, browse, and evaluate worlds as content-hashed bundles on Surface Area with the gatewaysdk benchmark hub client. - [Runs & Experiments](https://docs.surfacearea.ai/sdk/run-decorator.md): Track experiment runs, log metrics, group work into episodes, and post scores with the gatewaysdk run API. - [Sessions & Replay](https://docs.surfacearea.ai/sdk/sessions-replay.md): Load full sessions from Surface Area, dump them as test fixtures, and run regression suites from platform task sets. - [Hooks](https://docs.surfacearea.ai/sdk/hooks.md): Intercept an agent's tool calls from Python with rules in gateway-hooks.toml - deny, rewrite, add context or record a call from a command, an HTTP endpoint or a callable - and put a persona in front of every private individual with the built-in alias door. - [TypeScript SDK](https://docs.surfacearea.ai/sdk-ts/index.md): Install @withgateway/sdk, set credentials, and open a live world session from a Node agent. The import map for every subpath, and which extra packages each one needs. - [Worlds](https://docs.surfacearea.ai/sdk-ts/worlds.md): Resolve a world to a pinned version from Node, read its scenarios, dispatch a hosted run of a run config, and wait for the graded result. - [Describe a world and look things up by plaintext](https://docs.surfacearea.ai/sdk-ts/describe.md): Learn what a world holds and how each field is stored, then find a redacted row from the plaintext you know, with the same digest rule the runtime uses. - [Tool calls into rows](https://docs.surfacearea.ai/sdk-ts/ingest.md): Record an agent's real tool calls with captureToolCalls, then put them into a world with ingest - through your own transform, a declared ingest.toml mapping, or the operation's projection - and the world's contract. Redaction optional, off by default. - [World sessions](https://docs.surfacearea.ai/sdk-ts/sessions.md): Open a live world session from Node, every WorldSession method with its signature, the session token as the credential your client already sends, inline tasks, and stored tasks. - [Run a task suite](https://docs.surfacearea.ai/sdk-ts/run-sessions.md): runSessions opens one warm container per task, runs your agent against each in parallel, grades them, and gates the mean reward with failUnder. - [Gate a pull request](https://docs.surfacearea.ai/sdk-ts/ci.md): validateWorld pushes the pull request's own world source, dispatches a run config against it, waits for every run, gates on the score, and returns a comment-ready markdown summary. - [Drive a world UI](https://docs.surfacearea.ai/sdk-ts/browser.md): Open a Playwright browser on a world session's dashboard, hand the browser to a model as computer-use tools, and have every action recorded on the session timeline. - [Publish a world](https://docs.surfacearea.ai/sdk-ts/hub.md): Push a world's source as a content-hashed version, pull one back down, propose a change to a world you do not own, and create branches and tags, all from Node. - [Hooks](https://docs.surfacearea.ai/sdk-ts/hooks.md): Intercept an agent's tool calls with rules in gateway-hooks.toml - deny, rewrite, add context or record a call from a command, an HTTP endpoint or a function - and put a persona in front of every private individual with the built-in alias door. - [Scores and signals](https://docs.surfacearea.ai/sdk-ts/scores.md): Post evaluation scores against a trace or session from Node, and emit the success signals that feed a Success Metric, with the same shape the Python SDK sends. - [Sessions and replay](https://docs.surfacearea.ai/sdk-ts/replay.md): Load a recorded session from Node, and turn a platform scenario set into a regression suite that replays recorded tool outputs against your agent. - [Tracing and experiments](https://docs.surfacearea.ai/sdk-ts/tracing.md): Initialize tracing in a Node agent, wrap an LLM client, group a run into one trace, version agent builds, assign A/B variants, and let Surface Area gate which tools an agent may call. - [SDK Reference](https://docs.surfacearea.ai/sdk-reference/index.md): Exact signatures, parameters, and examples for every public gatewaysdk API, organized by feature area. - [Tracing](https://docs.surfacearea.ai/sdk-reference/tracing.md): Exact signatures, parameters, and runnable examples for the gatewaysdk tracing API: init, the span decorators, sessions, identify, semantic-convention keys, the exporter, and the auto-instrumentation registry. - [Runs & Scoring](https://docs.surfacearea.ai/sdk-reference/runs-scoring.md): Exact signatures for the gatewaysdk run, scoring, success-signal, and A/B experiment APIs, covering every parameter, default, return value, and platform mapping. - [Sessions, Replay & Users](https://docs.surfacearea.ai/sdk-reference/sessions-replay.md): Exact signatures for the gatewaysdk.sessions, gatewaysdk.replay, and gatewaysdk.users modules, covering loaders, replay suites, triage results, and end-user profiles. - [Overview](https://docs.surfacearea.ai/rest-api/index.md): The Surface Area REST API — worlds and world sessions first, then traces, scores and the rest — covering base URL, HTTP Basic authentication, pagination, timestamps, error shape, and the ingestion and MCP endpoints. - [Core Resources](https://docs.surfacearea.ai/rest-api/resources.md): Route tables and worked curl examples for the Surface Area REST API, covering worlds, world sessions, stored tasks, traces, sessions, observations, scores, datasets, annotation queues, media, metrics, and models. - [The MCP server](https://docs.surfacearea.ai/mcp/index.md): Let any MCP-compatible agent build worlds, seed them, store the tasks they are asked, run models against them, and query the traces that come back. - [Set up the MCP server](https://docs.surfacearea.ai/mcp/setup.md): Install the Node or Python client, set credentials, wire the server into Claude Code, Claude Desktop or Cursor, or call the JSON-RPC endpoint directly. - [Tool reference](https://docs.surfacearea.ai/mcp/tools.md): Every tool the Surface Area MCP server exposes, grouped by feature: connectors and worlds, world data, stored tasks, world sandboxes, runs and performance, traces and context. This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.