Skip to main content
Exact-signature reference for the run and scoring surface exported from gatewaysdk: the run context manager, the Run and episode Rollout classes, the module-level log/score/success/finish shortcuts, and the A/B experiments namespace.
For a walkthrough with end-to-end examples, read Runs & Experiments first. Every symbol below is imported from the top-level gatewaysdk package (the .run module), except experiments and init_experiments, which are top-level package attributes.
Runs are backend-connected. They read GATEWAY_PROJECT_ID, GATEWAY_HOST, GATEWAY_PUBLIC_KEY, and GATEWAY_SECRET_KEY from the environment. See Connect to Surface Area for credential setup.

The run and scoring symbols

Import each symbol directly from gatewaysdk.
gatewaysdk.Rollout is a different class from the episode returned by run.rollout(). The top-level Rollout (from gatewaysdk.types) is trainer rollout metadata with no scoring methods. The scoring episode documented on Score from an episode is never imported directly: you always get it from run.rollout(...).

Start a run with the run context manager

run() opens a run, yields a Run, and finishes it on exit, completing normally or marking failed when the body raises.
Every keyword after name is forwarded to the Run constructor.

Start a run manually with init_run

init_run() constructs a Run, starts it on the backend, and returns it. Call finish() yourself: no context manager wraps it.

Run constructor parameters

run(), init_run(), experiment(), and Run(...) all accept the same constructor keywords.
A missing project ID, a failed create request, or an invalid backend response raises gatewaysdk.GatewayRunError. The SDK does not fabricate a local run or run the body as if tracking had started. On create, it also records the current git SHA, branch, dirty state, and a config hash for reproducibility.

Decorate a function as an experiment or episode

experiment() wraps a whole function in a run; episode() wraps a function in an episode inside the already-active run.
Both decorators support synchronous and asynchronous functions. episode() raises RuntimeError when no run is active.

Run methods and properties

A Run tracks metrics, spawns episodes, posts scores, and finishes with a durable final state.

Read-only properties

Log metrics with run.log()

Metrics are buffered and flushed roughly once per second, so they appear on the run page while the run is still going. Logging to a finished run logs a warning and no-ops.

Other logging methods

Metric is a dataclass with fields name: str, value: float, step: Optional[int] = None, unit: str = "", metric_type: str = "gauge", and labels: Dict[str, str] = {}.

Create an episode with run.rollout()

rollout() returns an episode Rollout (a context manager) that carries its own session ID. Traces produced inside the episode are grouped under that session, and episode-level score() and success() calls target it. See Score from an episode.

Score a session or trace with run.score()

Each score targets exactly one session or trace: the API enforces this XOR. When a session is available (an explicit session_id or an active episode), the session wins even if a trace is also present, and session-targeted scores need no trace. Only without a session does score() use trace_id or fall back to the current span; with neither target, the score is skipped with a warning. Scores post in a background thread, and finish() waits for delivery. Scoring on a finished run warns and no-ops.

Feed a Success Metric with run.success()

success() is a thin wrapper over score() that tags the score as a success signal so the platform can roll it into an agent’s Success Metric.
A boolean value records a 1/0 BOOLEAN score; a number records a NUMERIC score. The call adds gateway.success.signal = True and gateway.success.metricName (and gateway.agent.id when agent_id is set) to the score metadata. Target resolution is identical to score().

Finish a run with run.finish()

state is one of "completed", "failed", or "cancelled". Finishing stops the heartbeat, drains buffered metrics, waits for in-flight score threads, and sends the final state.
completed is confirmed only after the SDK drains every accepted metric, joins accepted score requests, and the backend accepts the final state. If a metric or score was lost, or the final request is rejected or cannot be confirmed, finish() sets the run state to failed and raises GatewayRunError. Explicit failed and cancelled finalization stays best-effort so a telemetry failure never masks the original body exception.

Propagate remote cancellation

set_on_cancel(callback) registers a zero-argument callback fired when a “Cancel Run” click in the dashboard reaches the run through its heartbeat. Trainers set it to an algorithm’s cancel hook so a UI cancel stops the running algorithm.

Score from an episode with rollout.score()

An episode Rollout (returned by run.rollout()) captures its own session ID, and captures the first trace ID that starts inside its context. Its score() and success() methods auto-fill both, then delegate to the parent run.
Every keyword accepted by Run.score() and Run.success() passes through. The episode supplies session_id from itself and trace_id from its first span unless you override them.

Module-level shortcuts post to the active run

Four module-level functions act on whichever run is currently active, so shared helper code can log or score without threading a Run through every call.
log, score, and success raise RuntimeError when no run is active: call them inside with gatewaysdk.run(...), under @experiment, or after init_run(). finish is the exception: it quietly does nothing when no run is active.

Reach the active run or episode

Both return None when nothing is active, which makes logging optional in shared code.

Map score values to platform concepts

Every score() and success() call creates a Gateway Score entity. The data_type decides how the platform reads the value. success() infers BOOLEAN from a bool and NUMERIC from a number unless you set data_type. A boolean signal drives a Pass rate Success Metric; a numeric signal drives an Average, Count, or Sum. The agent_id argument sets gateway.agent.id on the score. Per-agent cost on the Agents page and per-agent Success Metric roll-ups key off the same ID, so a success() call tagged with agent_id reports next to that agent’s spend.
The signal name and the Success Metric name are separate. Set metric_name when they differ so an unrelated same-named score is never counted toward the metric.

Register and read A/B experiments

The experiments namespace assigns users to variants deterministically and logs exposures. Access it through the gatewaysdk.experiments proxy; no client instance is required.

Initialize the backend connection

init_experiments() starts the background exposure logger and creates the global experiment client. Call it once at startup for backend-synced experiments; skip it to run local-only (in-memory, no exposure logging).

The experiments namespace methods

register() takes each variant as a dict with key, weight, and config. get_variant() returns a Variant dataclass with key: str, weight: int, and config: Dict[str, Any].
When the backend is configured, get_variant() uses server-side balanced assignment for exact proportional allocation and falls back to local hash-based assignment if the backend is unreachable. The assigned variant is written to the OpenTelemetry span context, so traces produced afterward carry the experiment and variant.