Skip to main content
A Run groups one experiment’s work so you can watch metrics live and attach scores to the sessions or traces it produces. The run API has two equivalent styles: a context manager and a decorator. Both auto-start a run, send a heartbeat, and finish cleanly on exit or crash.
Runs require GATEWAY_PROJECT_ID. They also read GATEWAY_HOST, GATEWAY_PUBLIC_KEY, and GATEWAY_SECRET_KEY from the environment.
Runs are backend-connected. Missing credentials, a failed create request, or an invalid backend response raises gatewaysdk.GatewayRunError; the SDK does not fabricate a local Run or execute the context/decorated body as if tracking had started.

Start a run with a context manager

Use the context manager: it finishes the run even when the body raises. Pass a name and any run metadata to record.
gatewaysdk.run() yields a Run. The config, labels, and tags arguments are stored with the run for reproducibility and filtering.

Start a run with a decorator

Use @gatewaysdk.experiment when an entire function is the experiment. The run starts before the function runs and finishes after it returns or raises.
By default the Run is injected as the first argument. Set inject_run=False to omit it and reach the run through gatewaysdk.get_current_run() instead.
The decorator works on async functions too: it awaits the wrapped coroutine inside the run.

Log metrics

Call log() to record metrics. Pass a dict of name-value pairs, or a single name and value.
Metrics are buffered and flushed for live updates, so they appear on the run page while the run is still going. Outside the run object, gatewaysdk.log(metrics, step=...) logs to whichever run is currently active and raises RuntimeError if none is.

Group work into episodes

An episode (a Rollout) is one interaction sequence inside a run, such as a single task attempt. Each episode gets its own session id, so the traces it produces are grouped on the Sessions dashboard.
Create episodes with run.rollout(name) as a context manager. The decorator form is @gatewaysdk.episode, which requires an active run created by @experiment or with gatewaysdk.run(...).
Pass name_fn to build the episode name from the call arguments, or name for a static one. Set inject_rollout=True to receive the Rollout as the first argument.

Post scores to sessions or traces

score() attaches an evaluation result to exactly one session or trace, so it shows up on the Scores page and in session or trace detail. Scores differ from metrics: a metric is run-level, while a score grades a specific session or trace.
When you pass session_id or call score() while a Rollout is active, the session is preferred even if a trace is also available. Session-targeted scores work without a trace. Only when no session is available does score() use an explicit trace_id or fall back to the current OpenTelemetry span. If neither target is available, the score is skipped with a warning.
Each posted score targets exactly one session or trace. observation_id applies only to trace-targeted scores.
gatewaysdk.score(name, value) is the module-level shorthand that posts to the active run.

Feed a Success Metric

Use success() for a score that should feed a dashboard Success Metric. The signal name and metric identity are separate: name must match the metric’s configured SDK signal, while metric_name must match the Success Metric name. This prevents an unrelated same-named score or signal from being counted.
metric_name defaults to name, so you may omit it when the configured signal and Success Metric have the same name. The score follows the same target rules as score(): an explicit or active rollout session is preferred, then an explicit trace or the active OpenTelemetry trace.

Finish a run

The context manager and decorator finish the run for you. When you start a run manually with gatewaysdk.init_run(), call finish() yourself.
Valid final states are "completed", "failed", and "cancelled". gatewaysdk.finish(state) finishes the active run and does nothing if there is none. completed is confirmed only after the SDK drains every accepted metric, waits for accepted score requests, and the backend accepts the final state. If the SDK knows a metric or score was lost, or the final request is rejected or cannot be confirmed, it finishes local cleanup, sets the Run state to failed, and raises GatewayRunError instead of reporting success.
Explicit failed and cancelled finalization remains best-effort so a secondary telemetry failure does not replace the original body exception or cancellation.

Reach the active run or episode

Two helpers return the currently active objects from anywhere in your code.
They return None when no run or episode is active, which makes logging optional in shared code.

Symbol reference

  • Tracing shows how the spans that scores link to are produced.
  • Worlds client reads back the rollouts and per-task performance a run produces.
  • Evaluation & Replay covers how those scores are read against a pinned world version.