> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Sessions, Replay & Users

> Exact signatures for the gatewaysdk.sessions, gatewaysdk.replay, and gatewaysdk.users modules, covering loaders, replay suites, triage results, and end-user profiles.

Exact reference for three modules: `gatewaysdk.sessions` loads full sessions, `gatewaysdk.replay` runs regression suites from task sets, and `gatewaysdk.users` reads and writes end-user profiles.

<Info>
  For the narrative walkthrough (how to record a session, simulate the world, and triage results), read [Sessions & Replay](/sdk/sessions-replay) and [Agent Replay](/evaluation/agent-replay). Every signature below is verified against the SDK source.
</Info>

## Where each name lives

The package root re-exports the most common names, so `from gatewaysdk import load_session` and `from gatewaysdk.sessions import load_session` both work. Some names ship only on their submodule.

| Import | Re-exported at `gatewaysdk` root | Submodule-only |
| - | - | - |
| `gatewaysdk.sessions` | `load_session`, `SessionRecording`, `SessionTrace`, `SessionObservation`, `SessionLoadError` | none |
| `gatewaysdk.replay` | `ReplaySuite`, `ReplayCase`, `ReplayCheck`, `ToolStubRegistry`, `run_suite` | `suite`, `ReplayTask`, `ReplayReport`, `TraceDiff`, `DriftError`, `StubMissError`, `REPLAY_ENVIRONMENT` |
| `gatewaysdk.users` | `UsersClient`, `UserProfile`, `UserEvent`, `UserProfileError` | none |

## Load a session with `load_session`

Load one session (every trace, every observation, untruncated input and output) in a single HTTP call.

```python theme={null}
from gatewaysdk.sessions import load_session

def load_session(
    session_id: str,
    *,
    host: str | None = None,
    public_key: str | None = None,
    secret_key: str | None = None,
    timeout: float = 30.0,
) -> SessionRecording: ...
```

| Parameter | Type | Default | Description |
| - | - | - | - |
| `session_id` | `str` | required | The session to load |
| `host` | `str \| None` | `GATEWAY_HOST` | Platform base URL |
| `public_key` | `str \| None` | `GATEWAY_PUBLIC_KEY` | Project public key (`pk-lf-...`) |
| `secret_key` | `str \| None` | `GATEWAY_SECRET_KEY` | Project secret key (`sk-lf-...`) |
| `timeout` | `float` | `30.0` | Request timeout in seconds |

`load_session` always calls `GET /api/public/sessions/{session_id}?includeIO=true`, so a returned `SessionRecording` always carries full payloads for every trace and observation.

Credentials resolve in the same order the tracing exporter uses: the explicit keyword first, then the matching environment variable. A missing host or key raises `SessionLoadError` before any request. A `404` raises `SessionLoadError("Session not found: ...")`; any other non-`200` raises `SessionLoadError` with the status and response snippet.

## Inspect a `SessionRecording`

A fully loaded session. Traces arrive in API order; the selector methods re-sort chronologically by trace timestamp.

```python theme={null}
@dataclass
class SessionRecording:
    session_id: str
    project_id: str | None
    environment: str | None
    traces: list[SessionTrace]
```

| Method | Signature | Returns |
| - | - | - |
| `turns` | `turns()` | `list[SessionTrace]`: traces sorted chronologically, one per turn |
| `turn` | `turn(index: int)` | `SessionTrace`: the nth chronological turn (negative indices allowed) |
| `user_messages` | `user_messages()` | `list[str \| None]`: each turn's user message, in order |
| `final_output` | `final_output()` | `Any`: output of the last chronological turn (`None` when empty) |
| `tool_calls` | `tool_calls()` | `list[SessionObservation]`: every `TOOL` observation, chronological |
| `generations` | `generations()` | `list[SessionObservation]`: every `GENERATION` observation, chronological |
| `select` | `select(*, type=None, name=None, level=None)` | `list[SessionObservation]`: filtered observations across all turns |
| `to_dict` | `to_dict()` | `dict`: serializable form (keys camelCased: `sessionId`, `projectId`) |
| `to_fixture` | `to_fixture(path: str \| Path)` | `Path`: writes pretty JSON, creating parent directories |
| `from_api` | `from_api(raw: dict)` *(classmethod)* | `SessionRecording` |

`select` matches case-insensitively on `type` and `level`, and exactly on `name`; passing all three narrows further with logical AND.

### Dump a session to a fixture

`to_fixture` writes the full recording as indented JSON and returns the written path.

```python theme={null}
session = load_session("checkout-2f9c")
session.to_fixture("tests/fixtures/checkout_regression.json")
```

The file uses `to_dict`, so the on-disk shape matches what platform agents produce with their session-export tool. Human-written and agent-written fixtures are interchangeable.

## `SessionTrace`: one agent turn

```python theme={null}
@dataclass
class SessionTrace:
    id: str
    name: str | None
    timestamp: str | None
    input: Any
    output: Any
    metadata: Any = None
    tags: list[str] = []
    observations: list[SessionObservation] = []
```

| Member | Kind | Returns | Description |
| - | - | - | - |
| `tool_calls` | property | `list[SessionObservation]` | Observations whose `type` is `TOOL` |
| `generations` | property | `list[SessionObservation]` | Observations whose `type` is `GENERATION` |
| `user_message` | method | `str \| None` | The user message that started the turn |
| `from_api` | classmethod | `SessionTrace` | Builds a trace from raw API JSON |

`user_message()` handles three recorded input shapes: a plain string, an OpenAI-style message list (the last `role="user"` entry wins, so earlier entries stay as carried history), and a dict with a `content`, `input`, `text`, or `message` field.

## `SessionObservation`: one span

Every field is populated from the API with full input and output.

```python theme={null}
@dataclass
class SessionObservation:
    id: str
    type: str | None            # e.g. "GENERATION", "TOOL", "SPAN"
    name: str | None
    level: str | None
    status_message: str | None
    model: str | None
    input: Any
    output: Any
    metadata: Any = None
    parent_observation_id: str | None = None
    start_time: str | None = None
    end_time: str | None = None
```

`SessionObservation` exposes a `from_api(raw)` classmethod. `SessionLoadError` is a `RuntimeError` subclass raised for every session-load failure.

## Build a replay suite

A task set is a runnable regression corpus: each task carries a prompt, expected behavior, and provenance (`metadata.sourceSessionId`) back to the session that motivated it. `suite` and `ReplaySuite.from_task_set` load one into memory.

```python theme={null}
import gatewaysdk.replay as replay

# gatewaysdk.replay.suite: the shorthand entrypoint
def suite(name_or_id: str, **kwargs) -> ReplaySuite: ...

# ReplaySuite.from_task_set: the full classmethod (suite() forwards to it)
@classmethod
def from_task_set(
    cls,
    name_or_id: str,
    *,
    host: str | None = None,
    public_key: str | None = None,
    secret_key: str | None = None,
    timeout: float = 30.0,
) -> ReplaySuite: ...
```

| Parameter | Type | Default | Description |
| - | - | - | - |
| `name_or_id` | `str` | required | Task set name or `datasetId`, tried as an id first, then resolved by name |
| `host` | `str \| None` | `GATEWAY_HOST` | Platform base URL |
| `public_key` | `str \| None` | `GATEWAY_PUBLIC_KEY` | Project public key |
| `secret_key` | `str \| None` | `GATEWAY_SECRET_KEY` | Project secret key |
| `timeout` | `float` | `30.0` | Per-request timeout in seconds |

<Info>
  `suite` lives only on the submodule: call it as `gatewaysdk.replay.suite(...)`, not `gatewaysdk.suite(...)`. It accepts the same keyword arguments as `from_task_set` and forwards them unchanged.
</Info>

Loading pages tasks 200 at a time. A missing task set raises `ReplayError`; a non-`200` on the tasks or resolve call raises `ReplayError` with the status. The loaded suite retains the resolved credentials so a later `run` can record results back to the platform.

## `ReplaySuite` members

```python theme={null}
class ReplaySuite:
    task_set_id: str
    task_set_name: str
    cases: list[ReplayCase]
```

| Member | Signature | Returns | Description |
| - | - | - | - |
| `run` | `run(agent, *, miss_policy="error")` | `ReplayReport` | Run every case; delegates to `run_suite` |
| `snapshot` | `snapshot(directory: str \| Path)` | `list[Path]` | Pin each case's recorded session as `<case>.recording.json` |
| `__iter__` | `for case in suite` | `Iterator[ReplayCase]` | Iterate cases in task order |
| `__len__` | `len(suite)` | `int` | Case count |
| `__getitem__` | `suite[key]` | `ReplayCase` | Look up a case by `id` or `name` (raises `KeyError` if absent) |

<Info>
  `ReplaySuite.run` accepts only `agent` and `miss_policy`. To set `run_name` or opt out of recording, call the module function `run_suite(suite, agent, run_name=..., record=False)` directly; `run` does not forward those arguments.
</Info>

## Run a suite with `run_suite`

`run_suite` executes every case, triages each result, and (when the suite carries credentials) records the run back to the platform.

```python theme={null}
from gatewaysdk.replay import run_suite

def run_suite(
    suite: ReplaySuite,
    agent: Callable[..., Any],
    *,
    miss_policy: str = "error",
    record: bool | None = None,
    run_name: str | None = None,
) -> ReplayReport: ...
```

| Parameter | Type | Default | Description |
| - | - | - | - |
| `suite` | `ReplaySuite` | required | The loaded suite to run |
| `agent` | `Callable` | required | The agent under test (see signatures below) |
| `miss_policy` | `str` | `"error"` | Stub-miss behavior: `"error"`, `"reuse"`, or `"passthrough"` |
| `record` | `bool \| None` | `None` | `None` auto-enables recording when the suite was loaded from the platform; `False` disables it |
| `run_name` | `str \| None` | `None` | Names the recorded run; defaults to `replay-<8 hex chars>` |

The runner inspects the agent's parameter names and calls whichever shape it declares, most specific first.

| Declared signature | Runner calls |
| - | - |
| `agent(case)` | `agent(case)`: full `ReplayCase` |
| `agent(prompt, tools=...)` or `**kwargs` | `agent(case.prompt, tools=registry)` |
| `agent(prompt)` | `agent(case.prompt)` |

## Grade one case with `ReplayCase`

Each case binds a task to the session it was filed from. The recorded session loads lazily on first access to `recording`.

```python theme={null}
class ReplayCase:
    id: str            # property → task id
    name: str          # property → task name, falling back to id
    prompt: Any        # property → task prompt
    expected: Any      # property → task expected_output
    recording: SessionRecording   # property, lazy-loaded
```

| Method | Signature | Returns |
| - | - | - |
| `stubs` | `stubs(*, miss_policy="error", fallback=None, local=None)` | `ToolStubRegistry` built from the recorded session |
| `conversation` | `conversation()` | `list[str]`: the recorded per-turn user messages (falls back to `[str(prompt)]`) |
| `check` | `check(output=None, *, called_tools=None, stubs=None, mode=None)` | `ReplayCheck` |

Accessing `recording` when the task has no `sourceSessionId` raises `DriftError`: the case cannot be replayed without a source session.

### `check` parameters and modes

```python theme={null}
def check(
    self,
    output: Any = None,
    *,
    called_tools: list[str] | None = None,
    stubs: ToolStubRegistry | None = None,
    mode: str | None = None,
) -> ReplayCheck: ...
```

| Parameter | Type | Default | Description |
| - | - | - | - |
| `output` | `Any` | `None` | The agent's output (used for `outputExact` grading) |
| `called_tools` | `list[str] \| None` | `None` | Tool names the agent called; falls back to `stubs.called_names()` |
| `stubs` | `ToolStubRegistry \| None` | `None` | The registry used for the run; supplies drift and call log |
| `mode` | `str \| None` | `None` | `"contract"` (default) or `"exact"`; falls back to task `metadata.replayMode` |

Contract mode grades three invariants read from task metadata: `expectTools` (every listed tool must have been called), `mustNotCall` (none may have been called), and `outputExact: true` (the agent output must exactly equal `expected_output`). Exact mode adds a strict 1:1 trace comparison: the agent must reproduce every recorded tool call, in order, with nothing extra; any divergence is a `FAIL` and the full diff lands in `ReplayCheck.trace_diff`.

The verdict follows a fixed precedence: a contract violation always wins as `FAIL`, even on a drifted world; an unexplained omission is `FAIL` only on a clean world (drift downgrades it to `DRIFT`); a clean world with no violation is `PASS`.

## `ReplayCheck`: one case verdict

```python theme={null}
@dataclass
class ReplayCheck:
    status: str                    # "PASS" | "FAIL" | "DRIFT"
    failures: list[str] = []
    drift: list[str] = []
    trace_diff: TraceDiff | None = None
```

| Member | Kind | Returns | Description |
| - | - | - | - |
| `passed` | property | `bool` | `status == "PASS"` |
| `drifted` | property | `bool` | `status == "DRIFT"` |
| `explain` | method | `str` | Human-readable status plus per-line failure and drift reasons |

## `ReplayReport`: whole-suite result

`run` and `run_suite` return one report. Per-case verdicts are keyed by case name in `results`.

```python theme={null}
@dataclass
class ReplayReport:
    results: dict[str, ReplayCheck]
    run_name: str | None = None
    recording_errors: list[str] = []
```

| Member | Kind | Returns | Description |
| - | - | - | - |
| `passed` | property | `int` | Count of `PASS` cases |
| `failed` | property | `int` | Count of `FAIL` cases |
| `drifted` | property | `int` | Count of `DRIFT` cases |
| `ok` | property | `bool` | `failed == 0`: drift is a stale-test signal, not a failure |
| `explain` | method | `str` | Summary line plus non-passing detail and any recording errors |

```python theme={null}
report = replay.suite("checkout-regressions").run(my_agent)
assert report.ok                       # no FAILs (drift does not block)
report.results["earbuds-regression"]   # the ReplayCheck for one case
report.run_name                         # the platform run name, when recorded
report.recording_errors                 # non-empty if platform recording failed
```

## Simulate the world with `ToolStubRegistry`

The registry replays recorded tool outputs so an agent runs against the world exactly as it was recorded.

```python theme={null}
class ToolStubRegistry:
    def __init__(
        self,
        calls: list[RecordedToolCall],
        *,
        miss_policy: str = "error",
        fallback: Callable[[str, Any], Any] | None = None,
        local: dict[str, Callable[[Any], Any]] | None = None,
    ): ...
```

| Member | Signature | Returns |
| - | - | - |
| `from_session` | `from_session(session, *, miss_policy="error", fallback=None, local=None)` *(classmethod)* | `ToolStubRegistry` |
| `lookup` | `lookup(tool_name: str, tool_input=None)` | `Any`: the recorded output |
| `as_registry` | `as_registry(fallback=None, local=None)` | `StubRegistryMapping`: a `dict` of `name → callable`, drop-in for most tool layers |
| `called_names` | `called_names()` | `list[str]`: tool names in call order |
| `compare` | `compare(*, include_local=False)` | `TraceDiff`: replayed-vs-recorded 1:1 diff |
| `drifted` | property | `bool`: whether any drift event was recorded |

### Match strategies

`lookup` tries strategies in order and records the winning strategy in the call log.

| Strategy | When it applies |
| - | - |
| `local` | The tool was declared in `local=`; its live implementation runs, logged but never counted as drift |
| `exact` | Normalized-JSON equality of the tool input against an un-consumed recording |
| `exact-repeat` | The same `(name, input)` was already replayed; the recorded pair is served again |
| `ordinal` | The next un-consumed recording for that tool name (robust to argument phrasing drift) |

Auto-instrumented single-key envelopes (`args_json`, `arguments`, `kwargs`, `args`, `input`) are unwrapped on both sides before comparison, so recorded inputs match the arguments a live agent passes.

### Miss policy

When no strategy matches, `miss_policy` decides what happens. Every miss is recorded as drift regardless.

| `miss_policy` | Behavior |
| - | - |
| `"error"` | Raise `StubMissError` (strict replay) |
| `"reuse"` | Replay the last recorded output for that name again; never-recorded names still raise (or hit `fallback`) |
| `"passthrough"` | Call the `fallback` callable and record the miss as drift |

## `TraceDiff`, `ReplayTask`, and replay errors

```python theme={null}
@dataclass
class TraceDiff:
    matched: list[str] = []
    missing: list[str] = []
    extra: list[str] = []
    out_of_order: list[str] = []
```

`TraceDiff` exposes an `identical` property (`True` when nothing is missing, extra, or out of order) and an `explain()` method. It is the value on `ReplayCheck.trace_diff` and the return of `ToolStubRegistry.compare`.

```python theme={null}
@dataclass
class ReplayTask:
    id: str
    name: str | None
    prompt: Any
    expected_output: Any
    metadata: dict[str, Any] = {}
```

`ReplayTask` carries `source_session_id` and `source_trace_id` properties (read from `metadata.sourceSessionId` / `metadata.sourceTraceId`) and a `from_api` classmethod.

| Error | Base | Raised when |
| - | - | - |
| `ReplayError` | `RuntimeError` | Base class for suite loading and config problems |
| `DriftError` | `ReplayError` | The recorded session no longer matches the system: re-baseline the case |
| `StubMissError` | `KeyError` | A strict-mode (`"error"`) stub lookup found no recorded answer |

## Records written back to the platform

When recording is active, `run_suite` traces each case and files a dataset-run item so the run appears on the task set.

The one named module constant is `REPLAY_ENVIRONMENT`. The service name, trace name, session id, and tag scheme are applied inline in `run_suite`.

| Constant / convention | Value | Where it lands |
| - | - | - |
| `REPLAY_ENVIRONMENT` | `"gateway-replay"` | The environment each replay trace records under, hidden from the sessions and traces tables by default |
| service name | `"gatewaysdk-replay"` | The tracing service initialized for the run |
| trace name | `replay:<case name>` | Each case's replay trace |
| session id | `<run name>-<case name>` | Groups each case's replay as its own session |
| tags | `replay`, `taskSet:<name>`, `case:<name>`, `run:<run name>` | Applied to every replay trace |
| default run name | `replay-<8 hex chars>` | Used when `run_name` is not passed |

Each case is then filed as a dataset-run item via `POST /api/public/dataset-run-items`, carrying `runName`, `datasetItemId` (the case id), `traceId`, and a `metadata` object with `replay: true`, the `status`, `failures`, `drift`, and `sourceSessionId`. Because the replay trace reaches storage asynchronously, the recording call retries on `404`. Recording never raises: failures collect in `ReplayReport.recording_errors`.

## Pytest integration: `replay_cases`

The `gatewaysdk.replay.pytest` submodule parametrizes a test function with one case per task-set item.

```python theme={null}
from gatewaysdk.replay.pytest import replay_cases

@replay_cases("checkout-regressions")
def test_regression(case):
    stubs = case.stubs()
    out = my_agent(case.prompt, tools=stubs.as_registry())
    result = case.check(out, stubs=stubs)
    assert result.passed or result.drifted, result.explain()
```

```python theme={null}
def replay_cases(name_or_id: str, **suite_kwargs) -> Callable: ...
```

`replay_cases` forwards `**suite_kwargs` to `ReplaySuite.from_task_set`. When credentials are missing or the task set is unavailable, the cases **skip** (with the reason as the skip message) rather than fail, so an unconfigured local run stays green. An empty task set produces a single skipped `empty-task-set` parameter.

## Read and write profiles with `UsersClient`

`UsersClient` reads and upserts end-user profiles over the platform's Basic-auth endpoints.

```python theme={null}
from gatewaysdk import UsersClient

class UsersClient:
    def __init__(
        self,
        host: str,
        public_key: str,
        secret_key: str,
        *,
        timeout_seconds: int = 15,
    ): ...

    @classmethod
    def from_env(
        cls,
        *,
        host: str | None = None,
        public_key: str | None = None,
        secret_key: str | None = None,
        timeout_seconds: int = 15,
    ) -> UsersClient: ...
```

`from_env` reads `GATEWAY_HOST`, `GATEWAY_PUBLIC_KEY`, and `GATEWAY_SECRET_KEY`, using any explicit keyword first. A missing value raises `UserProfileError` naming the missing fields.

The client exposes four API methods: two for profiles and two for value events.

| Method | Signature | Returns | Endpoint |
| - | - | - | - |
| `get` | `get(user_id: str)` | `UserProfile` | `GET /api/public/user-profiles/{user_id}` |
| `update` | `update(user_id, *, display_name=None, email=None, attributes=None, notes=None, merge=True)` | `UserProfile` | `POST /api/public/user-profiles` |
| `track_event` | `track_event(user_id, event, *, value=None, metadata=None, trace_id=None, session_id=None)` | `UserEvent` | `POST /api/public/user-events` |
| `list_events` | `list_events(*, user_id=None, event=None, page=1, limit=50)` | `list[UserEvent]` | `GET /api/public/user-events` |

| `update` parameter | Type | Default | Description |
| - | - | - | - |
| `user_id` | `str` | required | The profile to upsert |
| `display_name` | `str \| None` | `None` | Sent only when not `None` |
| `email` | `str \| None` | `None` | Sent only when not `None` |
| `attributes` | `dict[str, str] \| None` | `None` | String key-value facts |
| `notes` | `str \| None` | `None` | Sent only when not `None` |
| `merge` | `bool` | `True` | `True` merges `attributes` with the stored set; `False` replaces it |

```python theme={null}
users = UsersClient.from_env()
users.update(
    "user-123",
    display_name="Acme Corp - Jane D.",
    attributes={"plan": "enterprise", "account_value": "48000"},
)
profile = users.get("user-123")
```

Any HTTP or network failure raises `UserProfileError` (a `RuntimeError` subclass) with the method, path, and response detail.

## `UserProfile`

The stored profile returned by `get` and `update`.

```python theme={null}
@dataclass
class UserProfile:
    user_id: str
    display_name: str | None
    email: str | None
    attributes: dict[str, str] = {}
    notes: str | None = None
```

## `UserEvent`

A recorded value event: one entry in the per-user ROI stream. `value` is the emitter-defined worth of the event (e.g. minutes saved, revenue); Surface Area sums it in the ROI and hours-saved rollups.

```python theme={null}
@dataclass
class UserEvent:
    id: str
    user_id: str
    event: str
    value: float | None
    metadata: dict[str, Any] = {}
    trace_id: str | None = None
    session_id: str | None = None
    created_at: str | None = None
```

```python theme={null}
users.track_event(
    "user-123",
    "ticket_resolved",
    value=12,                 # minutes of support time saved
    session_id=session_id,    # link to the run that produced it
)
recent = users.list_events(user_id="user-123", event="ticket_resolved")
```

## How `UsersClient` relates to `identify`

`gatewaysdk.identify` is the tracing-side counterpart to `UsersClient.update`. Both write the same profile; they differ in what else they do and when.

```python theme={null}
import gatewaysdk

gatewaysdk.identify(
    "user-123",
    attributes={"plan": "enterprise", "account_value": "48000"},
)
```

`identify` attributes the current trace to the user immediately (it stamps the open span and sets the user for every following span), then, when `attributes`, `display_name`, or `email` are given, updates the profile in a background thread via `UsersClient.from_env().update`. Use `identify` inside a traced request to attribute usage and enrich the profile in one call; use `UsersClient` directly for standalone profile reads (`get`) or synchronous writes outside a trace.

To attribute a session to an end user while it is being traced, see [Trace metadata](/tracing/metadata).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.