Skip to main content
Surface Area names six things: world, scenario, rollout, run, eval, and gate. Sessions seed a world, a scenario is attempted in it, a rollout is one attempt, a run is a batch of rollouts, an eval scores them, and a gate reads the scores. Each one uses the one before it, so read them in order the first time.

World

A sealed replay of the systems your agent touches, with its own tools, data and rules. A world is built from real sessions, so it behaves like the system your agent met in production. Nothing inside it reaches the outside, so a candidate agent can run against the same situation as many times as you like. Worlds live under Worlds in the project sidebar.

Scenario

One situation to attempt in a world, with the conditions that decide whether it passed. A scenario says what the agent is asked to do and what counts as doing it. A world plus a scenario is a repeatable test. Scenarios live under Scenarios. Your organization also ships ready-made benchmark suites of hundreds of scenarios each, so a project with no traffic yet still has something to run.

Rollout

One scenario, run once and graded. A rollout has a transcript, a result, and the scores the evals gave it.

Run

One model against one world version, holding every rollout it produced. The world version and the model are both pinned, so two runs differ by exactly the thing you changed and their pass rates can be read against each other.
A rollout is one attempt. A run is the batch. When a page shows you a pass rate, it is a run’s rate over its rollouts.

Eval

A judge that scores what an agent produced, so a run can be read as a pass rate. An eval reads a rollout, or a session or trace from production, and returns a score. Evals are written as LLM judges, as deterministic checks, or as humans reviewing in the queue. All three produce scores the same way. Evals live under Evals. Scoring your own production traces is covered in Evaluation & Replay.

Gate

The rule a run must clear to be promoted: enough of its graded rollouts pass, and there are enough of them for the rate to mean anything. A gate is two thresholds, not one. A 100 percent pass rate over three rollouts does not clear a gate that asks for fifty, because three results cannot tell you a rate. Gates are configured on the release settings for a project, and their verdicts show under Releases.

Words we do not use

These six replaced older words, which are gone from the product’s surfaces. URLs and API fields keep some of the historical names so existing links and integrations keep working. The words on screen are the ones above.