> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Runs & Experiments

> Track experiment runs, log metrics, group work into episodes, and post scores with the gatewaysdk run API.

A Run groups one experiment's work so you can watch metrics live and attach scores to the sessions or traces it produces. The run API has two equivalent styles: a context manager and a decorator. Both auto-start a run, send a heartbeat, and finish cleanly on exit or crash.

<Info>
  Runs require `GATEWAY_PROJECT_ID`. They also read `GATEWAY_HOST`, `GATEWAY_PUBLIC_KEY`, and `GATEWAY_SECRET_KEY` from the environment.
</Info>

<Info>
  Runs are backend-connected. Missing credentials, a failed create request, or an invalid backend response raises `gatewaysdk.GatewayRunError`; the SDK does not fabricate a local Run or execute the context/decorated body as if tracking had started.
</Info>

## Start a run with a context manager

Use the context manager: it finishes the run even when the body raises. Pass a name and any run metadata to record.

```python theme={null}
import gatewaysdk

with gatewaysdk.run("gepa-v1", config={"lr": 0.01}, tags=["search"]) as r:
    for epoch in range(100):
        loss = train_one_epoch()
        r.log({"loss": loss}, step=epoch)
```

`gatewaysdk.run()` yields a `Run`. The `config`, `labels`, and `tags` arguments are stored with the run for reproducibility and filtering.

## Start a run with a decorator

Use `@gatewaysdk.experiment` when an entire function is the experiment. The run starts before the function runs and finishes after it returns or raises.

```python theme={null}
import gatewaysdk

@gatewaysdk.experiment("gepa-v1", config={"lr": 0.01})
def train(run):
    run.log({"started": 1})
    for candidate in population:
        evaluate(candidate)

train()
```

By default the `Run` is injected as the first argument. Set `inject_run=False` to omit it and reach the run through `gatewaysdk.get_current_run()` instead.

```python theme={null}
@gatewaysdk.experiment(inject_run=False)
def train():
    gatewaysdk.log({"started": 1})
```

The decorator works on async functions too: it awaits the wrapped coroutine inside the run.

## Log metrics

Call `log()` to record metrics. Pass a dict of name-value pairs, or a single name and value.

```python theme={null}
with gatewaysdk.run("eval") as r:
    r.log({"accuracy": 0.92, "f1": 0.88}, step=10)
    r.log("latency_ms", 240.0, step=10, unit="ms")
```

Metrics are buffered and flushed for live updates, so they appear on the run page while the run is still going.

| Parameter | Type | Default | Description |
| - | - | - | - |
| `metrics` | `dict` or `str` | required | A `{name: value}` dict, or a single metric name |
| `value` | `float` | `None` | Required only when `metrics` is a single name |
| `step` | `int` | `None` | Training step or iteration number |
| `labels` | `dict` | `None` | Per-metric labels for filtering |
| `unit` | `str` | `""` | Unit of measurement, for example `"ms"` or `"tokens"` |
| `metric_type` | `str` | `"gauge"` | Visualization hint: `"gauge"`, `"counter"`, or `"histogram"` |

Outside the run object, `gatewaysdk.log(metrics, step=...)` logs to whichever run is currently active and raises `RuntimeError` if none is.

## Group work into episodes

An episode (a `Rollout`) is one interaction sequence inside a run, such as a single task attempt. Each episode gets its own session id, so the traces it produces are grouped on the Sessions dashboard.

```python theme={null}
with gatewaysdk.run("gepa-v1") as r:
    for candidate in population:
        with r.rollout(f"candidate-{candidate.id}") as episode:
            result = agent.execute(task)
            episode.score("rubric", grade(result))
```

Create episodes with `run.rollout(name)` as a context manager. The decorator form is `@gatewaysdk.episode`, which requires an active run created by `@experiment` or `with gatewaysdk.run(...)`.

```python theme={null}
@gatewaysdk.experiment("gepa")
def train(run):
    for candidate in population:
        evaluate(candidate)

@gatewaysdk.episode(name_fn=lambda c: f"candidate-{c.id}")
def evaluate(candidate):
    result = agent.execute(task)
```

Pass `name_fn` to build the episode name from the call arguments, or `name` for a static one. Set `inject_rollout=True` to receive the `Rollout` as the first argument.

## Post scores to sessions or traces

`score()` attaches an evaluation result to exactly one session or trace, so it shows up on the Scores page and in session or trace detail. Scores differ from metrics: a metric is run-level, while a score grades a specific session or trace.

```python theme={null}
with gatewaysdk.run("eval") as r:
    with r.rollout("task-1") as episode:
        result = agent.execute(task)
        episode.score("rubric", 0.85)        # targets the session; no trace required
        r.score("accuracy", 0.95, trace_id="abc123")  # or pass IDs explicitly
```

When you pass `session_id` or call `score()` while a Rollout is active, the session is preferred even if a trace is also available. Session-targeted scores work without a trace. Only when no session is available does `score()` use an explicit `trace_id` or fall back to the current OpenTelemetry span. If neither target is available, the score is skipped with a warning.

| Parameter | Type | Default | Description |
| - | - | - | - |
| `name` | `str` | required | Score name, for example `"rubric"` or `"accuracy"` |
| `value` | `float` or `str` | required | Numeric for `NUMERIC`, string for `CATEGORICAL` |
| `trace_id` | `str` | `None` | Trace to target when no session is available; auto-detected from the active span |
| `session_id` | `str` | `None` | Session to target; an explicit id or active episode takes precedence over a trace |
| `observation_id` | `str` | `None` | Span id for a trace-targeted score; omitted from session-targeted scores |
| `data_type` | `str` | `"NUMERIC"` | `"NUMERIC"`, `"CATEGORICAL"`, or `"BOOLEAN"` |
| `comment` | `str` | `None` | Optional explanation of the score |
| `metadata` | `dict` | `None` | Optional metadata |

<Info>
  Each posted score targets exactly one session or trace. `observation_id` applies only to trace-targeted scores.
</Info>

`gatewaysdk.score(name, value)` is the module-level shorthand that posts to the active run.

## Feed a Success Metric

Use `success()` for a score that should feed a dashboard Success Metric. The
signal name and metric identity are separate: `name` must match the metric's
configured SDK signal, while `metric_name` must match the Success Metric name.
This prevents an unrelated same-named score or signal from being counted.

```python theme={null}
with gatewaysdk.run("support-eval") as r:
    with r.rollout("ticket-42") as episode:
        resolved = agent.resolve(ticket)
        episode.success(
            "resolved_ticket",
            resolved,
            agent_id="support-triage",
            metric_name="Resolution rate",
        )
```

`metric_name` defaults to `name`, so you may omit it when the configured signal
and Success Metric have the same name. The score follows the same target rules
as `score()`: an explicit or active rollout session is preferred, then an
explicit trace or the active OpenTelemetry trace.

## Finish a run

The context manager and decorator finish the run for you. When you start a run manually with `gatewaysdk.init_run()`, call `finish()` yourself.

```python theme={null}
import gatewaysdk

r = gatewaysdk.init_run("manual-run", config={"seed": 7})
r.log({"step": 1})
r.finish(state="completed")
```

Valid final states are `"completed"`, `"failed"`, and `"cancelled"`. `gatewaysdk.finish(state)` finishes the active run and does nothing if there is none.

`completed` is confirmed only after the SDK drains every accepted metric, waits
for accepted score requests, and the backend accepts the final state. If the SDK
knows a metric or score was lost, or the final request is rejected or cannot be
confirmed, it finishes local cleanup, sets the Run state to `failed`, and raises
`GatewayRunError` instead of reporting success.

```python theme={null}
try:
    r.finish(state="completed")
except gatewaysdk.GatewayRunError as error:
    handle_delivery_failure(error)
```

Explicit `failed` and `cancelled` finalization remains best-effort so a secondary
telemetry failure does not replace the original body exception or cancellation.

## Reach the active run or episode

Two helpers return the currently active objects from anywhere in your code.

```python theme={null}
import gatewaysdk

run = gatewaysdk.get_current_run()        # Run or None
episode = gatewaysdk.get_current_rollout()  # Rollout or None
```

They return `None` when no run or episode is active, which makes logging optional in shared code.

## Symbol reference

| Symbol | Kind | Purpose |
| - | - | - |
| `run` | context manager | Start and finish a run around a block |
| `init_run` | function | Start a run you finish yourself |
| `experiment` | decorator | Wrap a function in a run |
| `episode` | decorator | Wrap a function in an episode within a run |
| `log` | function | Log metrics to the active run |
| `score` | function | Post a score from the active run |
| `finish` | function | Finish the active run |
| `get_current_run` | function | Return the active `Run` or `None` |
| `get_current_rollout` | function | Return the active `Rollout` or `None` |
| `GatewayRunError` | exception | Run creation or successful completion could not meet its durability contract |
| `Run` | class | A run; `log`, `score`, `rollout`, `finish` |
| `Rollout` | class | An episode; `score`, `session_id`, `trace_id` |
| `Metric` | dataclass | A metric with full metadata for `log_metrics` |

## Related pages

* [Tracing](/tracing) shows how the spans that scores link to are produced.
* [Worlds client](/sdk/environments) reads back the rollouts and per-task performance a run produces.
* [Evaluation & Replay](/evaluation) covers how those scores are read against a pinned world version.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.