Runs require
GATEWAY_PROJECT_ID. They also read GATEWAY_HOST, GATEWAY_PUBLIC_KEY, and GATEWAY_SECRET_KEY from the environment.Runs are backend-connected. Missing credentials, a failed create request, or an invalid backend response raises
gatewaysdk.GatewayRunError; the SDK does not fabricate a local Run or execute the context/decorated body as if tracking had started.Start a run with a context manager
Use the context manager: it finishes the run even when the body raises. Pass a name and any run metadata to record.gatewaysdk.run() yields a Run. The config, labels, and tags arguments are stored with the run for reproducibility and filtering.
Start a run with a decorator
Use@gatewaysdk.experiment when an entire function is the experiment. The run starts before the function runs and finishes after it returns or raises.
Run is injected as the first argument. Set inject_run=False to omit it and reach the run through gatewaysdk.get_current_run() instead.
Log metrics
Calllog() to record metrics. Pass a dict of name-value pairs, or a single name and value.
Outside the run object,
gatewaysdk.log(metrics, step=...) logs to whichever run is currently active and raises RuntimeError if none is.
Group work into episodes
An episode (aRollout) is one interaction sequence inside a run, such as a single task attempt. Each episode gets its own session id, so the traces it produces are grouped on the Sessions dashboard.
run.rollout(name) as a context manager. The decorator form is @gatewaysdk.episode, which requires an active run created by @experiment or with gatewaysdk.run(...).
name_fn to build the episode name from the call arguments, or name for a static one. Set inject_rollout=True to receive the Rollout as the first argument.
Post scores to sessions or traces
score() attaches an evaluation result to exactly one session or trace, so it shows up on the Scores page and in session or trace detail. Scores differ from metrics: a metric is run-level, while a score grades a specific session or trace.
session_id or call score() while a Rollout is active, the session is preferred even if a trace is also available. Session-targeted scores work without a trace. Only when no session is available does score() use an explicit trace_id or fall back to the current OpenTelemetry span. If neither target is available, the score is skipped with a warning.
Each posted score targets exactly one session or trace.
observation_id applies only to trace-targeted scores.gatewaysdk.score(name, value) is the module-level shorthand that posts to the active run.
Feed a Success Metric
Usesuccess() for a score that should feed a dashboard Success Metric. The
signal name and metric identity are separate: name must match the metric’s
configured SDK signal, while metric_name must match the Success Metric name.
This prevents an unrelated same-named score or signal from being counted.
metric_name defaults to name, so you may omit it when the configured signal
and Success Metric have the same name. The score follows the same target rules
as score(): an explicit or active rollout session is preferred, then an
explicit trace or the active OpenTelemetry trace.
Finish a run
The context manager and decorator finish the run for you. When you start a run manually withgatewaysdk.init_run(), call finish() yourself.
"completed", "failed", and "cancelled". gatewaysdk.finish(state) finishes the active run and does nothing if there is none.
completed is confirmed only after the SDK drains every accepted metric, waits
for accepted score requests, and the backend accepts the final state. If the SDK
knows a metric or score was lost, or the final request is rejected or cannot be
confirmed, it finishes local cleanup, sets the Run state to failed, and raises
GatewayRunError instead of reporting success.
failed and cancelled finalization remains best-effort so a secondary
telemetry failure does not replace the original body exception or cancellation.
Reach the active run or episode
Two helpers return the currently active objects from anywhere in your code.None when no run or episode is active, which makes logging optional in shared code.
Symbol reference
Related pages
- Tracing shows how the spans that scores link to are produced.
- Worlds client reads back the rollouts and per-task performance a run produces.
- Evaluation & Replay covers how those scores are read against a pinned world version.