Skip to main content
Run agents against reality before reality matters. Surface Area rebuilds the systems your agent works against — your ledger, your service desk, your identity provider — as a sealed world. Your agent calls the world as often as it likes, and nothing it does reaches a live system. A candidate version replays the situations your production agent already met and comes back scored.
Get Started installs the command-line tool and walks you to a graded run inside a hosted world in about ten minutes.

What makes this different

Most tools in this space report what your agent did. Surface Area reports what a version you have not shipped yet will do. Deciding whether a change is safe takes somewhere to run it against the same conditions, as many times as it takes for the result to mean something.

The loop

1

Build a world

Start a world from a connector template, a connector your organization published, or a contract you write yourself. The platform hosts it, serves the vendor’s routes, and pins every version. See Worlds.
2

Give the world its material

Instrument your agent and Surface Area records every LLM call, tool call, and nested step as a trace. Those sessions become the rows and the situations a world is built from. See Tracing.
3

Ask it questions

One captured session becomes many scenarios. Withhold a piece of evidence, poison another, make a tool start failing, push the horizon out, or fork at the turn where it went wrong.
4

Run a candidate against all of them

Every rollout is graded by deterministic checks, LLM judges, or humans. The batch is a run. The world version and the model are both pinned, so two runs differ by exactly the thing you changed. See Evaluation & Replay.
5

Let a gate decide

A gate asks two questions before a build is promoted: did enough rollouts pass, and were there enough of them for the rate to mean anything.

Start here

Get Started

Worlds

Glossary

Evaluation & Replay

Tracing

Python SDK

Six words

The product is built out of six nouns: world, scenario, rollout, run, eval, and gate. Each one uses the one before it. Read the Glossary first.

What else is here

Worlds. Build one, put data in it, open a session against it, run an agent, and version it as your systems move. Evaluation & Replay. How a world’s runs are scored. Write the criteria a judge grades against, replay real sessions as regression tests, and keep a world’s scenarios in one place. Tracing. A few lines of Python capture each call as a trace over OpenTelemetry. Traces are useful on their own, and they are the raw material every world is built from. The gateway CLI. Worlds, sessions, data, and runs from a terminal or a continuous-integration job. Python SDK. The tracing decorator, the worlds client, runs, and replay, with a TypeScript SDK alongside.

For AI agents

A machine-readable index of this documentation is published at /docs/llms.txt for coding assistants and agents.