> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Agent Replay

> Turn real production sessions into a regression suite - replay them against a new agent version and read the pass, fail, and drift results as sessions and scored runs in Surface Area.

Agent Replay re-runs recorded production sessions against a new version of your agent. Curate the sessions worth guarding into a scenario set, replay them, and read the pass, fail, and drift results back as sessions and scored runs.

## Why replay keeps agent changes safe

Agent inputs are open-ended and agent tools touch live systems, so an ordinary unit test cannot reproduce the multi-turn conversation that triggered a bug.

A recorded session holds every turn, every tool call, and every output. Replaying it runs your new agent against the exact situation a user hit, with the recorded tool outputs standing in for the live systems. The result is a deterministic regression test built from real traffic.

<Info>
  Replay and [worlds](/worlds) answer two different questions. Replay tests the
  path that actually happened, with the recorded answers played back. A world
  tests the paths that could happen, because the systems underneath it hold state
  and respond to whatever the agent does. Replay costs less to set up; use a world
  when the behaviour depends on the system's own reaction.
</Info>

## The replay loop: capture, curate, replay, compare

Four steps take a real session from production to a release gate.

<Steps>
  <Step title="Capture real sessions">
    Your agent's runs are already recorded as sessions in [Tracing](/tracing), each with its full, untruncated payload. Give each session a stable, meaningful id at record time so you can find it later.
  </Step>

  <Step title="Curate the representative ones into a scenario set">
    Pick the sessions worth guarding -- a bug you just fixed, a tricky edge case, a golden path -- and file them into a [scenario set](./data-task-sets). Add each one by hand, or promote a failure straight from its session. Each scenario keeps a link back to the session it came from.
  </Step>

  <Step title="Replay them against a new agent version">
    Point the SDK at the scenario set and run your new agent over every case. The SDK loads each scenario's recorded session and feeds your agent the recorded tool outputs, so the run stays deterministic. Grade in **contract** mode (the agent must call the right tools and avoid the forbidden ones) or **exact** mode (every recorded call reproduced, in order). See [Sessions & Replay](/sdk/sessions-replay) for the code.
  </Step>

  <Step title="Compare the results">
    Each case comes back as a new session and a per-case verdict: **PASS**, **FAIL**, or **DRIFT**. Compare the run against the previous version to see what changed, then gate the release on the outcome.
  </Step>
</Steps>

<Info>
  A **DRIFT** verdict means the recording no longer matches your system -- a tool was renamed, or the agent asked for something the recording does not hold. Drift is a stale *test*, not a broken agent: re-record the session instead of blocking the change.
</Info>

## Reading replay results in the app

Every replay run records itself back to Surface Area as ordinary sessions. Each case's replay becomes its own session named `<run>-<case>`, tagged `replay`, `taskSet:<name>`, `case:<name>`, and `run:<name>`. The `taskSet` tag keeps the historical name for a scenario set, which is what you filter on.

*A replay session in detail view: the Session view tab lists the session's traces, with Annotate and Run Evaluator in the toolbar and a Run log tab alongside.*

Open a replay session to read its **Session view** -- the conversation turn by turn, with every score attached -- and switch to the **Run log** tab for the raw traces behind it.

Replay traffic lands in a dedicated `gateway-replay` environment, which Surface Area hides from the sessions and traces tables by default, so regression runs never inflate your real-usage views. Switch the environment filter to `gateway-replay` to see replay sessions in the tables, or open them from the scenario set's runs, where each replay run appears as a dataset run with one item and one verdict per case.

## Wiring replay into reviews and releases

For a verdict on a single case, open its replay session and use the same tools you use on live traffic: **Annotate** it yourself, add it to **Human Review** for a reviewer's rubric, or **Run Evaluator** to score the replay trace automatically. An [evaluator](/evaluation) grades a replay session the same way it grades any other.

For releases, the suite's headline result is your gate: the run is `ok` only when no case fails, so a continuous-integration job can block a merge on a single assertion. A [gate](/glossary#gate) reads the same scores when it decides whether a build is promoted.

## Set it up in code

Load a suite from a scenario set and run your agent over it:

```python theme={null}
import gatewaysdk.replay as replay

report = replay.suite("checkout-regressions").run(my_agent, run_name="pre-deploy-7")
assert report.ok   # ok == no FAILs; drift does not block
```

`replay.suite()` takes the name of a scenario set, which the SDK addresses by its `taskSet` identifier. The full reference -- recording named sessions, stubbing the recorded calls, contract versus exact grading, and the pytest integration -- is in [Sessions & Replay](/sdk/sessions-replay).

## Related pages

* [Sessions & Replay](/sdk/sessions-replay) is the SDK reference for loading sessions and running replay suites.
* [Scenarios & task sets](./data-task-sets) is where you curate the sessions each replay runs against.
* [Evaluation (LLM-as-Judge)](/evaluation) scores a replay trace the same way it scores live runs.
* [Worlds](/worlds) is where you go when the recorded answers are not enough and the systems underneath have to react.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.