> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Sessions & Replay

> Load full sessions from Surface Area, dump them as test fixtures, and run regression suites from platform task sets.

Turn real production sessions into regression tests. The SDK can load any
session with **full, untruncated payloads** (every trace, every observation)
and replay it against your agent with the recorded tool outputs standing in
for the live world.

Replay checks whether the agent still does what it did on the recorded path.
When the outcome depends on how the surrounding systems behave, build a
[world](/worlds) from the recorded session and run the agent against the world
instead of a transcript.

## Record a named session

Give a session a stable id at record time. That id is what you file into a
task set and replay against later:

```python theme={null}
import gatewaysdk.tracing as tracing

with tracing.session("checkout-regression-001"):
    with tracing.trace("turn-1", input=prompt):
        ...   # every trace, LLM call, and tool span in here is grouped
```

### One trace per turn

A session should read turn-by-turn: **one trace per user message**. With
auto-instrumented frameworks each top-level agent invocation becomes one
trace, so run one invocation per turn and carry the history yourself. Don't
push a whole conversation through a single call, or the session collapses
into one giant trace:

```python theme={null}
with tracing.session("checkout-regression-001"):
    history = None
    for message in user_messages:
        turn_input = [*history, {"role": "user", "content": message}] if history else message
        result = await Runner.run(agent, turn_input)     # one trace per turn
        history = result.to_input_list()
```

Recorded this way, `session.turns()` maps 1:1 to the conversation, which
turn-level selectors and full-conversation replay depend on.

## Load a session

```python theme={null}
from gatewaysdk.sessions import load_session

session = load_session("session-id")   # uses GATEWAY_HOST / _PUBLIC_KEY / _SECRET_KEY
```

`load_session` calls `GET /api/public/sessions/{id}?includeIO=true`: one
request returns the complete session. Credentials resolve from the same
environment variables tracing uses, or pass `host=`, `public_key=`,
`secret_key=` explicitly.

### Work with the recording

```python theme={null}
for turn in session.turns():           # traces in chronological order
    print(turn.input, "→", turn.output)

session.turn(0)                        # nth turn (negative indices ok)
session.turn(0).user_message()         # the user message that started the turn
session.user_messages()                # every turn's user message, in order
session.final_output()                 # last turn's output
session.tool_calls()                   # every TOOL observation, session-wide
session.generations()                  # every GENERATION observation
session.select(type="TOOL", name="checkout")   # generic filter
```

Each observation carries `input`, `output`, `model`, `level`,
`status_message`, and timestamps, untruncated.

### Dump a fixture

```python theme={null}
session.to_fixture("tests/fixtures/checkout_regression.json")
```

Writes the whole session as pretty JSON: the same shape platform agents
produce with their `export_session` tool, so human-made and agent-made
fixtures are interchangeable.

## Replay suites from task sets

A **task set** on Surface Area is a regression corpus: each task stores a
prompt, the expected behavior, and provenance to the **recorded session** that
motivated it (`metadata.sourceSessionId`). Sessions get filed into task sets
by hand or from a failing rollout (`promote_failure_to_scenario`): no repo change needed to add a case.

```python theme={null}
from gatewaysdk.replay import ReplaySuite

suite = ReplaySuite.from_task_set("checkout-regressions")   # name or datasetId

for case in suite:
    print(case.name, "→", case.prompt)
    recording = case.recording    # full SessionRecording, loaded lazily
```

### Simulate the recorded world

The source session recorded every tool output, so replays are deterministic:
your agent runs against the catalog/database/API responses *as they were*:

```python theme={null}
case = suite["earbuds-regression"]
stubs = case.stubs()                          # ToolStubRegistry

out = my_agent(case.prompt, tools=stubs.as_registry())
```

Stub matching tries strategies in order, and it understands
auto-instrumented recordings (OpenInference `{"args_json": "..."}` wrappers
and agent-as-tool `{"input": "..."}` envelopes are unwrapped before
comparison, so recorded inputs match the args your live agent passes):

1. **local**: tools you declared live (see below); never drift
2. **exact**: normalized-JSON input equality
3. **exact-repeat**: identical re-reads of a `(name, input)` pair return the
   same recording again instead of exhausting it
4. **ordinal**: nth un-consumed recording for that name (robust to arg
   phrasing drift across models/prompts)

Misses follow `miss_policy`:

* `"error"` (default): raise `StubMissError`, strict replay
* `"reuse"`: recordings for that tool exhausted? Replay the last one again
  so research-heavy agents (more searches than the recording holds) finish
  the run. Never-recorded names still raise.
* `"passthrough"`: call your `fallback=lambda name, args: ...` live

Either way every miss is recorded as **drift**.

Tools your recording predates (or deterministic helpers that never need
recording) run live with `local=`, visible to the contract check but never
counted as drift:

```python theme={null}
stubs = case.stubs(miss_policy="reuse", local={"compare_products": my_impl})
```

Most agent frameworks (openai-agents included) catch tool exceptions and hand
them to the model as error output, so the `"error"` policy does not hard-stop
those runs: a dead recording usually surfaces as a max-turns exception
instead. Catch both, then triage with `case.check`.

### Grade with three-way triage

```python theme={null}
result = case.check(out, stubs=stubs)
assert result.passed, result.explain()
```

| Status | Meaning | What to do |
| - | - | - |
| `PASS` | Contract held in a successfully simulated world | Nothing |
| `FAIL` | The agent violated the contract (`expectTools` missing, `mustNotCall` hit) | A real regression: block the change |
| `DRIFT` | The recorded world no longer matches the system (stub miss, renamed tool) | The **test** is stale, not the agent: re-record the session |

Contract invariants come from task metadata: `expectTools` (must all be
called), `mustNotCall` (none may be called), and optional
`outputExact: true` for literal output matching. Semantic grading stays with
you: assert on `case.expected`, or score the replay trace with a platform
evaluator.

### Grade against the ENTIRE session: 1:1 trace replay

Every case retains its **full recorded session** (`case.recording`): all
turns, all observations, untruncated, not just the prompt and expected
output. Two things build on that:

**Re-drive the whole conversation.** `case.conversation()` returns the
recorded user message of each turn (each turn was one trace); feed them to
your agent in order to replay the full session, not just turn one.

**Exact mode.** For a true 1:1 comparison, grade with `mode="exact"` (or set
`metadata.replayMode: "exact"` on the task):

```python theme={null}
result = case.check(final_reply, stubs=stubs, mode="exact")
print(result.trace_diff.explain())
# 1:1 replay: all 9 recorded tool calls reproduced in order
#   -- or --
# trace diverges from recording (3 matched):
#   not reproduced: deal_scout("best wireless earbuds")
#   extra call: search_products({"query": "usb chargers"}) [reuse]
#   out of order: add_to_cart({"product": "Flux 65W GaN Charger"})
```

Any divergence (a recorded call not reproduced, an extra call, or the same
calls in a different order) is a FAIL, with the full diff in
`result.trace_diff`. Use it for deterministic/pinned agents (temperature 0,
workflow-style agents); for stochastic agents prefer contract mode, and run
exact mode to see how behavior shifted. Agents making parallel tool calls can
interleave order legitimately: the recorded order is the observation order of
the original trace.

You can also assert on anything in the recording directly:

```python theme={null}
recording = case.recording
recording.turns()               # every turn (trace), in order
recording.turn(1).tool_calls    # a specific turn's tools, with full IO
recording.final_output()        # recorded final answer, e.g. for a judge
recording.user_messages()       # per-turn user inputs
```

### Run the whole suite

```python theme={null}
import gatewaysdk.replay as replay

report = replay.suite("checkout-regressions").run(my_agent)
print(report.explain())                  # "14 PASS · 1 FAIL · 1 DRIFT" + details
assert report.ok                         # ok == no FAILs (drift doesn't block)
report.run_name                          # the platform run this was recorded as
report.recording_errors                  # non-empty if platform recording failed
```

Suite runs are **recorded back to the platform** automatically when the
suite was loaded with credentials (opt out with `record=False`):

* each case's replay becomes its own **session** named `<run>-<case>`, traced
  under the internal `gateway-replay` environment and tagged `replay`,
  `taskSet:<name>`, `case:<name>`, `run:<runName>`, hidden from the sessions
  page by default; filter on the `taskSet:<name>` tag to find them
* the run is filed on the task set as a dataset run (one item per case, with
  the PASS/FAIL/DRIFT status and the source session id in item metadata),
  listed on the **Runs** tab behind **Edit simulations**, so you can compare
  runs across agent versions and score them with evaluators

Any live SDK traffic you tag `taskSet:<name>` yourself is found by the same
tag filter: task sets work as test suites for non-environment rollouts too.

Your agent can declare any of these signatures; the runner adapts:

```python theme={null}
def my_agent(prompt, tools): ...   # the common shape (tools = stub registry)
def my_agent(prompt): ...          # prompt-only agents
def my_agent(case): ...            # full ReplayCase (case.stubs(), case.recording)
```

### Pytest integration

```python theme={null}
import pytest
from gatewaysdk.replay.pytest import replay_cases

@replay_cases("checkout-regressions")
def test_regression(case):
    stubs = case.stubs(miss_policy="reuse")
    out = my_agent(case.prompt, tools=stubs.as_registry())
    result = case.check(out, stubs=stubs)
    assert result.status != "FAIL", result.explain()   # regression → block
    if result.drifted:                                 # stale recording →
        pytest.xfail(result.explain())                 # flag, don't block
```

One parametrized test per task-set case. When Surface Area credentials aren't
configured the cases **skip** instead of failing, so local runs without keys
stay green.

### Pin the corpus into your repo

```python theme={null}
suite.snapshot("tests/fixtures/")   # one <case>.recording.json per case
```

## Task set API

Manage suites programmatically (Basic auth with project API keys):

| Endpoint | Purpose |
| - | - |
| `GET /api/public/task-sets?name=` | List / resolve task sets |
| `POST /api/public/task-sets` | Create a task set |
| `GET /api/public/task-sets/{datasetId}/tasks` | List tasks (prompt, expectedOutput, metadata incl. `sourceSessionId`, and `author`: the userId or `apikey:{id}` that filed it) |
| `POST /api/public/task-sets/{datasetId}/tasks` | File a task: set `metadata.sourceSessionId` to link the session it came from |

Task metadata drives grading: `expectTools`, `mustNotCall`, `outputExact`,
and `replayMode: "exact"` for 1:1 grading, all optional.

## Where to go next

* [Worlds client](/sdk/environments) — run the agent against a live world instead of a recording.
* [Evaluation & Replay](/evaluation) — the platform side of replay and scoring.
* [Tracing sessions](/tracing/sessions) — how a session gets recorded in the first place.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.