Skip to main content
Exact reference for three modules: gatewaysdk.sessions loads full sessions, gatewaysdk.replay runs regression suites from task sets, and gatewaysdk.users reads and writes end-user profiles.
For the narrative walkthrough (how to record a session, simulate the world, and triage results), read Sessions & Replay and Agent Replay. Every signature below is verified against the SDK source.

Where each name lives

The package root re-exports the most common names, so from gatewaysdk import load_session and from gatewaysdk.sessions import load_session both work. Some names ship only on their submodule.

Load a session with load_session

Load one session (every trace, every observation, untruncated input and output) in a single HTTP call.
load_session always calls GET /api/public/sessions/{session_id}?includeIO=true, so a returned SessionRecording always carries full payloads for every trace and observation. Credentials resolve in the same order the tracing exporter uses: the explicit keyword first, then the matching environment variable. A missing host or key raises SessionLoadError before any request. A 404 raises SessionLoadError("Session not found: ..."); any other non-200 raises SessionLoadError with the status and response snippet.

Inspect a SessionRecording

A fully loaded session. Traces arrive in API order; the selector methods re-sort chronologically by trace timestamp.
select matches case-insensitively on type and level, and exactly on name; passing all three narrows further with logical AND.

Dump a session to a fixture

to_fixture writes the full recording as indented JSON and returns the written path.
The file uses to_dict, so the on-disk shape matches what platform agents produce with their session-export tool. Human-written and agent-written fixtures are interchangeable.

SessionTrace: one agent turn

user_message() handles three recorded input shapes: a plain string, an OpenAI-style message list (the last role="user" entry wins, so earlier entries stay as carried history), and a dict with a content, input, text, or message field.

SessionObservation: one span

Every field is populated from the API with full input and output.
SessionObservation exposes a from_api(raw) classmethod. SessionLoadError is a RuntimeError subclass raised for every session-load failure.

Build a replay suite

A task set is a runnable regression corpus: each task carries a prompt, expected behavior, and provenance (metadata.sourceSessionId) back to the session that motivated it. suite and ReplaySuite.from_task_set load one into memory.
suite lives only on the submodule: call it as gatewaysdk.replay.suite(...), not gatewaysdk.suite(...). It accepts the same keyword arguments as from_task_set and forwards them unchanged.
Loading pages tasks 200 at a time. A missing task set raises ReplayError; a non-200 on the tasks or resolve call raises ReplayError with the status. The loaded suite retains the resolved credentials so a later run can record results back to the platform.

ReplaySuite members

ReplaySuite.run accepts only agent and miss_policy. To set run_name or opt out of recording, call the module function run_suite(suite, agent, run_name=..., record=False) directly; run does not forward those arguments.

Run a suite with run_suite

run_suite executes every case, triages each result, and (when the suite carries credentials) records the run back to the platform.
The runner inspects the agent’s parameter names and calls whichever shape it declares, most specific first.

Grade one case with ReplayCase

Each case binds a task to the session it was filed from. The recorded session loads lazily on first access to recording.
Accessing recording when the task has no sourceSessionId raises DriftError: the case cannot be replayed without a source session.

check parameters and modes

Contract mode grades three invariants read from task metadata: expectTools (every listed tool must have been called), mustNotCall (none may have been called), and outputExact: true (the agent output must exactly equal expected_output). Exact mode adds a strict 1:1 trace comparison: the agent must reproduce every recorded tool call, in order, with nothing extra; any divergence is a FAIL and the full diff lands in ReplayCheck.trace_diff. The verdict follows a fixed precedence: a contract violation always wins as FAIL, even on a drifted world; an unexplained omission is FAIL only on a clean world (drift downgrades it to DRIFT); a clean world with no violation is PASS.

ReplayCheck: one case verdict

ReplayReport: whole-suite result

run and run_suite return one report. Per-case verdicts are keyed by case name in results.

Simulate the world with ToolStubRegistry

The registry replays recorded tool outputs so an agent runs against the world exactly as it was recorded.

Match strategies

lookup tries strategies in order and records the winning strategy in the call log. Auto-instrumented single-key envelopes (args_json, arguments, kwargs, args, input) are unwrapped on both sides before comparison, so recorded inputs match the arguments a live agent passes.

Miss policy

When no strategy matches, miss_policy decides what happens. Every miss is recorded as drift regardless.

TraceDiff, ReplayTask, and replay errors

TraceDiff exposes an identical property (True when nothing is missing, extra, or out of order) and an explain() method. It is the value on ReplayCheck.trace_diff and the return of ToolStubRegistry.compare.
ReplayTask carries source_session_id and source_trace_id properties (read from metadata.sourceSessionId / metadata.sourceTraceId) and a from_api classmethod.

Records written back to the platform

When recording is active, run_suite traces each case and files a dataset-run item so the run appears on the task set. The one named module constant is REPLAY_ENVIRONMENT. The service name, trace name, session id, and tag scheme are applied inline in run_suite. Each case is then filed as a dataset-run item via POST /api/public/dataset-run-items, carrying runName, datasetItemId (the case id), traceId, and a metadata object with replay: true, the status, failures, drift, and sourceSessionId. Because the replay trace reaches storage asynchronously, the recording call retries on 404. Recording never raises: failures collect in ReplayReport.recording_errors.

Pytest integration: replay_cases

The gatewaysdk.replay.pytest submodule parametrizes a test function with one case per task-set item.
replay_cases forwards **suite_kwargs to ReplaySuite.from_task_set. When credentials are missing or the task set is unavailable, the cases skip (with the reason as the skip message) rather than fail, so an unconfigured local run stays green. An empty task set produces a single skipped empty-task-set parameter.

Read and write profiles with UsersClient

UsersClient reads and upserts end-user profiles over the platform’s Basic-auth endpoints.
from_env reads GATEWAY_HOST, GATEWAY_PUBLIC_KEY, and GATEWAY_SECRET_KEY, using any explicit keyword first. A missing value raises UserProfileError naming the missing fields. The client exposes four API methods: two for profiles and two for value events.
Any HTTP or network failure raises UserProfileError (a RuntimeError subclass) with the method, path, and response detail.

UserProfile

The stored profile returned by get and update.

UserEvent

A recorded value event: one entry in the per-user ROI stream. value is the emitter-defined worth of the event (e.g. minutes saved, revenue); Surface Area sums it in the ROI and hours-saved rollups.

How UsersClient relates to identify

gatewaysdk.identify is the tracing-side counterpart to UsersClient.update. Both write the same profile; they differ in what else they do and when.
identify attributes the current trace to the user immediately (it stamps the open span and sets the user for every following span), then, when attributes, display_name, or email are given, updates the profile in a background thread via UsersClient.from_env().update. Use identify inside a traced request to attribute usage and enrich the profile in one call; use UsersClient directly for standalone profile reads (get) or synchronous writes outside a trace. To attribute a session to an end user while it is being traced, see Trace metadata.