Skip to main content
Surface Area lets you test an agent against a sealed copy of the systems it works with, before it touches the real ones. You build a world that behaves like a vendor’s API and data, give it scenarios to attempt, run your agent through them, and read a graded pass rate you can compare from one version to the next.

How the pieces fit

1

Build a world

A sealed, hosted copy of a system: its data model, routes, tools and rows. Start from a shipped vendor template, an OpenAPI spec, your own API, or a contract you write.
2

Add scenarios

Each scenario is a task for the agent plus the checks that decide whether it passed.
3

Run your agent

Open a session per scenario and point your agent at the world’s URL and token instead of the vendor’s. Nothing it does leaves the world.
4

Grade and compare

Every attempt is graded. A run’s pass rate is pinned to one world version and one model, so two runs differ only by what you changed.
5

Gate a release

A gate turns the pass rate into a merge or ship decision, in CI or on the Releases page.
The Glossary defines these words in one page.

What the platform offers

Ways to drive it

Pick your starting point