@withgateway/sdk/sessions loads one. @withgateway/sdk/replay turns a set of them into a regression suite that runs your current agent against the world as it was.
Neither subpath needs OpenTelemetry. Both read the same three environment variables as everything else.
Load one session
loadSession(sessionId, options?) fetches every trace and observation of a session, with untruncated inputs and outputs.
options accepts host, publicKey, secretKey and timeoutSeconds, which defaults to 30. Missing credentials and a missing session both throw SessionLoadError.
A
SessionTrace carries the same toolCalls and generations for one turn, plus userMessage().
Pin a session as a fixture
toFixture(path) writes the recording to a file. Commit it and a test suite replays without reaching the platform at all.
Replay a scenario set as a regression suite
A scenario set on the platform is a runnable suite. Each case carries its prompt, what was expected, and a link to the recorded session that motivated it.suite(nameOrId, options?) is shorthand for ReplaySuite.fromTaskSet. runSuite(suite, agent, options?) is the same thing as suite.run(agent, options?) in function form.
Three layers pin three different things
- Tool stubs pin the world: recorded tool outputs make a replay deterministic.
- Invariants pin the contract: the tools that must and must not be called.
- Your own assertions pin the semantics.
Triage is three-way
report.ok is true when nothing failed. Drift is stale-test signal, not a failure.
A forbidden tool call fails even on a drifted world, because a call that happened is positive evidence the contract broke. An expected tool that was never called only fails on a clean world, because drift already explains the omission.
The report
Each
ReplayCheck carries status, failures, drift, traceDiff and its own explain().
Drive the loop yourself
Iterating the suite gives you each case, its stubs and its check. Use it when your agent does not fit the(prompt, tools) shape.
testCase exposes id, name, prompt, expected, task and recording(), which loads the session the case was filed from.
Stub options
stubs(options?) builds the recorded world for one case.
reuse replays the nearest recorded answer again, which keeps a run going when the agent explores more than the recording covers. Each reuse is reported as drift.
Runner options
Replay traffic records under the
gateway-replay environment, which the platform hides from the sessions and traces tables by default.
Snapshot a whole suite into the repository
snapshot(directory) writes every case’s recorded session as a fixture, and returns the paths it wrote. Cases with no source session are skipped.
Where to go next
- Evaluation and Replay for how scenario sets are built.
- Sessions, Replay and Users for the Python equivalents.
- Scores and signals for recording what a replay concluded.