gatewaysdk.sessions loads full sessions, gatewaysdk.replay runs regression suites from task sets, and gatewaysdk.users reads and writes end-user profiles.
For the narrative walkthrough (how to record a session, simulate the world, and triage results), read Sessions & Replay and Agent Replay. Every signature below is verified against the SDK source.
Where each name lives
The package root re-exports the most common names, sofrom gatewaysdk import load_session and from gatewaysdk.sessions import load_session both work. Some names ship only on their submodule.
Load a session with load_session
Load one session (every trace, every observation, untruncated input and output) in a single HTTP call.
load_session always calls GET /api/public/sessions/{session_id}?includeIO=true, so a returned SessionRecording always carries full payloads for every trace and observation.
Credentials resolve in the same order the tracing exporter uses: the explicit keyword first, then the matching environment variable. A missing host or key raises SessionLoadError before any request. A 404 raises SessionLoadError("Session not found: ..."); any other non-200 raises SessionLoadError with the status and response snippet.
Inspect a SessionRecording
A fully loaded session. Traces arrive in API order; the selector methods re-sort chronologically by trace timestamp.
select matches case-insensitively on type and level, and exactly on name; passing all three narrows further with logical AND.
Dump a session to a fixture
to_fixture writes the full recording as indented JSON and returns the written path.
to_dict, so the on-disk shape matches what platform agents produce with their session-export tool. Human-written and agent-written fixtures are interchangeable.
SessionTrace: one agent turn
user_message() handles three recorded input shapes: a plain string, an OpenAI-style message list (the last role="user" entry wins, so earlier entries stay as carried history), and a dict with a content, input, text, or message field.
SessionObservation: one span
Every field is populated from the API with full input and output.
SessionObservation exposes a from_api(raw) classmethod. SessionLoadError is a RuntimeError subclass raised for every session-load failure.
Build a replay suite
A task set is a runnable regression corpus: each task carries a prompt, expected behavior, and provenance (metadata.sourceSessionId) back to the session that motivated it. suite and ReplaySuite.from_task_set load one into memory.
suite lives only on the submodule: call it as gatewaysdk.replay.suite(...), not gatewaysdk.suite(...). It accepts the same keyword arguments as from_task_set and forwards them unchanged.ReplayError; a non-200 on the tasks or resolve call raises ReplayError with the status. The loaded suite retains the resolved credentials so a later run can record results back to the platform.
ReplaySuite members
ReplaySuite.run accepts only agent and miss_policy. To set run_name or opt out of recording, call the module function run_suite(suite, agent, run_name=..., record=False) directly; run does not forward those arguments.Run a suite with run_suite
run_suite executes every case, triages each result, and (when the suite carries credentials) records the run back to the platform.
The runner inspects the agent’s parameter names and calls whichever shape it declares, most specific first.
Grade one case with ReplayCase
Each case binds a task to the session it was filed from. The recorded session loads lazily on first access to recording.
Accessing
recording when the task has no sourceSessionId raises DriftError: the case cannot be replayed without a source session.
check parameters and modes
Contract mode grades three invariants read from task metadata:
expectTools (every listed tool must have been called), mustNotCall (none may have been called), and outputExact: true (the agent output must exactly equal expected_output). Exact mode adds a strict 1:1 trace comparison: the agent must reproduce every recorded tool call, in order, with nothing extra; any divergence is a FAIL and the full diff lands in ReplayCheck.trace_diff.
The verdict follows a fixed precedence: a contract violation always wins as FAIL, even on a drifted world; an unexplained omission is FAIL only on a clean world (drift downgrades it to DRIFT); a clean world with no violation is PASS.
ReplayCheck: one case verdict
ReplayReport: whole-suite result
run and run_suite return one report. Per-case verdicts are keyed by case name in results.
Simulate the world with ToolStubRegistry
The registry replays recorded tool outputs so an agent runs against the world exactly as it was recorded.
Match strategies
lookup tries strategies in order and records the winning strategy in the call log.
Auto-instrumented single-key envelopes (
args_json, arguments, kwargs, args, input) are unwrapped on both sides before comparison, so recorded inputs match the arguments a live agent passes.
Miss policy
When no strategy matches,miss_policy decides what happens. Every miss is recorded as drift regardless.
TraceDiff, ReplayTask, and replay errors
TraceDiff exposes an identical property (True when nothing is missing, extra, or out of order) and an explain() method. It is the value on ReplayCheck.trace_diff and the return of ToolStubRegistry.compare.
ReplayTask carries source_session_id and source_trace_id properties (read from metadata.sourceSessionId / metadata.sourceTraceId) and a from_api classmethod.
Records written back to the platform
When recording is active,run_suite traces each case and files a dataset-run item so the run appears on the task set.
The one named module constant is REPLAY_ENVIRONMENT. The service name, trace name, session id, and tag scheme are applied inline in run_suite.
Each case is then filed as a dataset-run item via
POST /api/public/dataset-run-items, carrying runName, datasetItemId (the case id), traceId, and a metadata object with replay: true, the status, failures, drift, and sourceSessionId. Because the replay trace reaches storage asynchronously, the recording call retries on 404. Recording never raises: failures collect in ReplayReport.recording_errors.
Pytest integration: replay_cases
The gatewaysdk.replay.pytest submodule parametrizes a test function with one case per task-set item.
replay_cases forwards **suite_kwargs to ReplaySuite.from_task_set. When credentials are missing or the task set is unavailable, the cases skip (with the reason as the skip message) rather than fail, so an unconfigured local run stays green. An empty task set produces a single skipped empty-task-set parameter.
Read and write profiles with UsersClient
UsersClient reads and upserts end-user profiles over the platform’s Basic-auth endpoints.
from_env reads GATEWAY_HOST, GATEWAY_PUBLIC_KEY, and GATEWAY_SECRET_KEY, using any explicit keyword first. A missing value raises UserProfileError naming the missing fields.
The client exposes four API methods: two for profiles and two for value events.
UserProfileError (a RuntimeError subclass) with the method, path, and response detail.
UserProfile
The stored profile returned by get and update.
UserEvent
A recorded value event: one entry in the per-user ROI stream. value is the emitter-defined worth of the event (e.g. minutes saved, revenue); Surface Area sums it in the ROI and hours-saved rollups.
How UsersClient relates to identify
gatewaysdk.identify is the tracing-side counterpart to UsersClient.update. Both write the same profile; they differ in what else they do and when.
identify attributes the current trace to the user immediately (it stamps the open span and sets the user for every following span), then, when attributes, display_name, or email are given, updates the profile in a background thread via UsersClient.from_env().update. Use identify inside a traced request to attribute usage and enrich the profile in one call; use UsersClient directly for standalone profile reads (get) or synchronous writes outside a trace.
To attribute a session to an end user while it is being traced, see Trace metadata.