Skip to main content
Agent Replay re-runs recorded production sessions against a new version of your agent. Curate the sessions worth guarding into a scenario set, replay them, and read the pass, fail, and drift results back as sessions and scored runs.

Why replay keeps agent changes safe

Agent inputs are open-ended and agent tools touch live systems, so an ordinary unit test cannot reproduce the multi-turn conversation that triggered a bug. A recorded session holds every turn, every tool call, and every output. Replaying it runs your new agent against the exact situation a user hit, with the recorded tool outputs standing in for the live systems. The result is a deterministic regression test built from real traffic.
Replay and worlds answer two different questions. Replay tests the path that actually happened, with the recorded answers played back. A world tests the paths that could happen, because the systems underneath it hold state and respond to whatever the agent does. Replay costs less to set up; use a world when the behaviour depends on the system’s own reaction.

The replay loop: capture, curate, replay, compare

Four steps take a real session from production to a release gate.
1

Capture real sessions

Your agent’s runs are already recorded as sessions in Tracing, each with its full, untruncated payload. Give each session a stable, meaningful id at record time so you can find it later.
2

Curate the representative ones into a scenario set

Pick the sessions worth guarding — a bug you just fixed, a tricky edge case, a golden path — and file them into a scenario set. Add each one by hand, or promote a failure straight from its session. Each scenario keeps a link back to the session it came from.
3

Replay them against a new agent version

Point the SDK at the scenario set and run your new agent over every case. The SDK loads each scenario’s recorded session and feeds your agent the recorded tool outputs, so the run stays deterministic. Grade in contract mode (the agent must call the right tools and avoid the forbidden ones) or exact mode (every recorded call reproduced, in order). See Sessions & Replay for the code.
4

Compare the results

Each case comes back as a new session and a per-case verdict: PASS, FAIL, or DRIFT. Compare the run against the previous version to see what changed, then gate the release on the outcome.
A DRIFT verdict means the recording no longer matches your system — a tool was renamed, or the agent asked for something the recording does not hold. Drift is a stale test, not a broken agent: re-record the session instead of blocking the change.

Reading replay results in the app

Every replay run records itself back to Surface Area as ordinary sessions. Each case’s replay becomes its own session named <run>-<case>, tagged replay, taskSet:<name>, case:<name>, and run:<name>. The taskSet tag keeps the historical name for a scenario set, which is what you filter on. A replay session in detail view: the Session view tab lists the session’s traces, with Annotate and Run Evaluator in the toolbar and a Run log tab alongside. Open a replay session to read its Session view — the conversation turn by turn, with every score attached — and switch to the Run log tab for the raw traces behind it. Replay traffic lands in a dedicated gateway-replay environment, which Surface Area hides from the sessions and traces tables by default, so regression runs never inflate your real-usage views. Switch the environment filter to gateway-replay to see replay sessions in the tables, or open them from the scenario set’s runs, where each replay run appears as a dataset run with one item and one verdict per case.

Wiring replay into reviews and releases

For a verdict on a single case, open its replay session and use the same tools you use on live traffic: Annotate it yourself, add it to Human Review for a reviewer’s rubric, or Run Evaluator to score the replay trace automatically. An evaluator grades a replay session the same way it grades any other. For releases, the suite’s headline result is your gate: the run is ok only when no case fails, so a continuous-integration job can block a merge on a single assertion. A gate reads the same scores when it decides whether a build is promoted.

Set it up in code

Load a suite from a scenario set and run your agent over it:
replay.suite() takes the name of a scenario set, which the SDK addresses by its taskSet identifier. The full reference — recording named sessions, stubbing the recorded calls, contract versus exact grading, and the pytest integration — is in Sessions & Replay.
  • Sessions & Replay is the SDK reference for loading sessions and running replay suites.
  • Scenarios & task sets is where you curate the sessions each replay runs against.
  • Evaluation (LLM-as-Judge) scores a replay trace the same way it scores live runs.
  • Worlds is where you go when the recorded answers are not enough and the systems underneath have to react.