> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Sessions and replay

> Load a recorded session from Node, and turn a platform scenario set into a regression suite that replays recorded tool outputs against your agent.

A recorded session is what an agent actually did, with full inputs and outputs. `@withgateway/sdk/sessions` loads one. `@withgateway/sdk/replay` turns a set of them into a regression suite that runs your current agent against the world as it was.

Neither subpath needs OpenTelemetry. Both read the same three environment variables as everything else.

## Load one session

`loadSession(sessionId, options?)` fetches every trace and observation of a session, with untruncated inputs and outputs.

```typescript theme={null}
import { loadSession } from "@withgateway/sdk/sessions";

const recording = await loadSession("sess_9f21c4");

console.log(recording.sessionId, recording.environment);
console.log(recording.turns().length);
console.log(recording.userMessages());
console.log(recording.finalOutput());
```

`options` accepts `host`, `publicKey`, `secretKey` and `timeoutSeconds`, which defaults to 30. Missing credentials and a missing session both throw `SessionLoadError`.

| Method | Returns | What it gives you |
| - | - | - |
| `turns()` | `SessionTrace[]` | Every trace of the session, in order |
| `turn(index)` | `SessionTrace` | One turn, with negative indexing from the end |
| `userMessages()` | `(string \| null)[]` | The user input of each turn |
| `finalOutput()` | `unknown` | The last turn's output |
| `toolCalls()` | `SessionObservation[]` | Every tool call across the session |
| `generations()` | `SessionObservation[]` | Every model call across the session |
| `select(filter)` | `SessionObservation[]` | Observations matching `{ type?, name?, level? }` |
| `toJSON()` | `Record<string, unknown>` | The whole recording as plain data |
| `toFixture(path)` | `string` | Writes the recording to disk and returns the path |

A `SessionTrace` carries the same `toolCalls` and `generations` for one turn, plus `userMessage()`.

## Pin a session as a fixture

`toFixture(path)` writes the recording to a file. Commit it and a test suite replays without reaching the platform at all.

```typescript theme={null}
const recording = await loadSession("sess_9f21c4");
recording.toFixture("tests/fixtures/refund-flow.recording.json");
```

## Replay a scenario set as a regression suite

A scenario set on the platform is a runnable suite. Each case carries its prompt, what was expected, and a link to the recorded session that motivated it.

```typescript theme={null}
import { suite } from "@withgateway/sdk/replay";

const regressions = await suite("checkout-regressions");
const report = await regressions.run(async (prompt, tools) => {
  return myAgent(prompt, tools);
});

console.log(report.explain());
process.exit(report.ok ? 0 : 1);
```

`suite(nameOrId, options?)` is shorthand for `ReplaySuite.fromTaskSet`. `runSuite(suite, agent, options?)` is the same thing as `suite.run(agent, options?)` in function form.

### Three layers pin three different things

* **Tool stubs** pin the world: recorded tool outputs make a replay deterministic.
* **Invariants** pin the contract: the tools that must and must not be called.
* **Your own assertions** pin the semantics.

### Triage is three-way

| Status | Meaning | What to do |
| - | - | - |
| `PASS` | The contract held | Nothing |
| `FAIL` | A real regression: the world simulated fine and the contract broke | Fix the agent |
| `DRIFT` | The recorded world no longer matches the system, from stub misses or renamed tools | Re-record the session; the test is stale, not the agent |

`report.ok` is true when nothing failed. Drift is stale-test signal, not a failure.

<Info>
  A forbidden tool call fails even on a drifted world, because a call that happened is positive evidence the contract broke. An expected tool that was never called only fails on a clean world, because drift already explains the omission.
</Info>

## The report

| Member | Type | What it is |
| - | - | - |
| `results` | `Record<string, ReplayCheck>` | One entry per case, keyed by case name |
| `passed`, `failed`, `drifted` | `number` | Counts per status |
| `ok` | `boolean` | True when `failed` is zero |
| `runName` | `string \| null` | The run the cases were filed under, when recording |
| `recordingErrors` | `string[]` | Cases whose result could not be filed back |
| `explain()` | `string` | A printable summary of everything that was not a pass |

Each `ReplayCheck` carries `status`, `failures`, `drift`, `traceDiff` and its own `explain()`.

## Drive the loop yourself

Iterating the suite gives you each case, its stubs and its check. Use it when your agent does not fit the `(prompt, tools)` shape.

```typescript theme={null}
import assert from "node:assert";
import { suite } from "@withgateway/sdk/replay";

const regressions = await suite("checkout-regressions");

for (const testCase of regressions) {
  const stubs = await testCase.stubs({ missPolicy: "reuse" });
  const output = await myAgent(testCase.prompt, stubs.asRegistry());
  const result = testCase.check(output, { calledTools: stubs.calledNames() });
  assert.ok(result.passed, result.explain());
}
```

`testCase` exposes `id`, `name`, `prompt`, `expected`, `task` and `recording()`, which loads the session the case was filed from.

### Stub options

`stubs(options?)` builds the recorded world for one case.

| Option | Type | Meaning |
| - | - | - |
| `missPolicy` | `"error" \| "reuse" \| "passthrough"` | What happens when the agent calls something the recording does not hold; `error` is the default |
| `local` | `Record<string, LocalTool>` | Tools to run live instead of replaying, for new or deterministic tools |
| `fallback` | `(toolName, toolInput) => unknown` | Answer a miss yourself |

`reuse` replays the nearest recorded answer again, which keeps a run going when the agent explores more than the recording covers. Each reuse is reported as drift.

### Runner options

| Option | Type | Default | Meaning |
| - | - | - | - |
| `missPolicy` | `MissPolicy` | `"error"` | Passed through to every case's stubs |
| `record` | `boolean` | true when the suite was loaded from the platform | File each case as a run item on the scenario set |
| `runName` | `string` | a generated name | What the run is called on the platform |

Replay traffic records under the `gateway-replay` environment, which the platform hides from the sessions and traces tables by default.

## Snapshot a whole suite into the repository

`snapshot(directory)` writes every case's recorded session as a fixture, and returns the paths it wrote. Cases with no source session are skipped.

```typescript theme={null}
const paths = await regressions.snapshot("tests/fixtures");
```

## Where to go next

* [Evaluation and Replay](/evaluation) for how scenario sets are built.
* [Sessions, Replay and Users](/sdk-reference/sessions-replay) for the Python equivalents.
* [Scores and signals](/sdk-ts/scores) for recording what a replay concluded.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.