Skip to main content
Turn real production sessions into regression tests. The SDK can load any session with full, untruncated payloads (every trace, every observation) and replay it against your agent with the recorded tool outputs standing in for the live world. Replay checks whether the agent still does what it did on the recorded path. When the outcome depends on how the surrounding systems behave, build a world from the recorded session and run the agent against the world instead of a transcript.

Record a named session

Give a session a stable id at record time. That id is what you file into a task set and replay against later:

One trace per turn

A session should read turn-by-turn: one trace per user message. With auto-instrumented frameworks each top-level agent invocation becomes one trace, so run one invocation per turn and carry the history yourself. Don’t push a whole conversation through a single call, or the session collapses into one giant trace:
Recorded this way, session.turns() maps 1:1 to the conversation, which turn-level selectors and full-conversation replay depend on.

Load a session

load_session calls GET /api/public/sessions/{id}?includeIO=true: one request returns the complete session. Credentials resolve from the same environment variables tracing uses, or pass host=, public_key=, secret_key= explicitly.

Work with the recording

Each observation carries input, output, model, level, status_message, and timestamps, untruncated.

Dump a fixture

Writes the whole session as pretty JSON: the same shape platform agents produce with their export_session tool, so human-made and agent-made fixtures are interchangeable.

Replay suites from task sets

A task set on Surface Area is a regression corpus: each task stores a prompt, the expected behavior, and provenance to the recorded session that motivated it (metadata.sourceSessionId). Sessions get filed into task sets by hand or from a failing rollout (promote_failure_to_scenario): no repo change needed to add a case.

Simulate the recorded world

The source session recorded every tool output, so replays are deterministic: your agent runs against the catalog/database/API responses as they were:
Stub matching tries strategies in order, and it understands auto-instrumented recordings (OpenInference {"args_json": "..."} wrappers and agent-as-tool {"input": "..."} envelopes are unwrapped before comparison, so recorded inputs match the args your live agent passes):
  1. local: tools you declared live (see below); never drift
  2. exact: normalized-JSON input equality
  3. exact-repeat: identical re-reads of a (name, input) pair return the same recording again instead of exhausting it
  4. ordinal: nth un-consumed recording for that name (robust to arg phrasing drift across models/prompts)
Misses follow miss_policy:
  • "error" (default): raise StubMissError, strict replay
  • "reuse": recordings for that tool exhausted? Replay the last one again so research-heavy agents (more searches than the recording holds) finish the run. Never-recorded names still raise.
  • "passthrough": call your fallback=lambda name, args: ... live
Either way every miss is recorded as drift. Tools your recording predates (or deterministic helpers that never need recording) run live with local=, visible to the contract check but never counted as drift:
Most agent frameworks (openai-agents included) catch tool exceptions and hand them to the model as error output, so the "error" policy does not hard-stop those runs: a dead recording usually surfaces as a max-turns exception instead. Catch both, then triage with case.check.

Grade with three-way triage

Contract invariants come from task metadata: expectTools (must all be called), mustNotCall (none may be called), and optional outputExact: true for literal output matching. Semantic grading stays with you: assert on case.expected, or score the replay trace with a platform evaluator.

Grade against the ENTIRE session: 1:1 trace replay

Every case retains its full recorded session (case.recording): all turns, all observations, untruncated, not just the prompt and expected output. Two things build on that: Re-drive the whole conversation. case.conversation() returns the recorded user message of each turn (each turn was one trace); feed them to your agent in order to replay the full session, not just turn one. Exact mode. For a true 1:1 comparison, grade with mode="exact" (or set metadata.replayMode: "exact" on the task):
Any divergence (a recorded call not reproduced, an extra call, or the same calls in a different order) is a FAIL, with the full diff in result.trace_diff. Use it for deterministic/pinned agents (temperature 0, workflow-style agents); for stochastic agents prefer contract mode, and run exact mode to see how behavior shifted. Agents making parallel tool calls can interleave order legitimately: the recorded order is the observation order of the original trace. You can also assert on anything in the recording directly:

Run the whole suite

Suite runs are recorded back to the platform automatically when the suite was loaded with credentials (opt out with record=False):
  • each case’s replay becomes its own session named <run>-<case>, traced under the internal gateway-replay environment and tagged replay, taskSet:<name>, case:<name>, run:<runName>, hidden from the sessions page by default; filter on the taskSet:<name> tag to find them
  • the run is filed on the task set as a dataset run (one item per case, with the PASS/FAIL/DRIFT status and the source session id in item metadata), listed on the Runs tab behind Edit simulations, so you can compare runs across agent versions and score them with evaluators
Any live SDK traffic you tag taskSet:<name> yourself is found by the same tag filter: task sets work as test suites for non-environment rollouts too. Your agent can declare any of these signatures; the runner adapts:

Pytest integration

One parametrized test per task-set case. When Surface Area credentials aren’t configured the cases skip instead of failing, so local runs without keys stay green.

Pin the corpus into your repo

Task set API

Manage suites programmatically (Basic auth with project API keys): Task metadata drives grading: expectTools, mustNotCall, outputExact, and replayMode: "exact" for 1:1 grading, all optional.

Where to go next