> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Runs & Scoring

> Exact signatures for the gatewaysdk run, scoring, success-signal, and A/B experiment APIs, covering every parameter, default, return value, and platform mapping.

Exact-signature reference for the run and scoring surface exported from `gatewaysdk`: the `run` context manager, the `Run` and episode `Rollout` classes, the module-level `log`/`score`/`success`/`finish` shortcuts, and the A/B experiments namespace.

<Info>
  For a walkthrough with end-to-end examples, read [Runs & Experiments](/sdk/run-decorator) first. Every symbol below is imported from the top-level `gatewaysdk` package (the `.run` module), except `experiments` and `init_experiments`, which are top-level package attributes.
</Info>

Runs are backend-connected. They read `GATEWAY_PROJECT_ID`, `GATEWAY_HOST`, `GATEWAY_PUBLIC_KEY`, and `GATEWAY_SECRET_KEY` from the environment. See [Connect to Surface Area](/sdk#connect-to-surface-area) for credential setup.

## The run and scoring symbols

Import each symbol directly from `gatewaysdk`.

| Symbol | Kind | Purpose |
| - | - | - |
| `run` | context manager | Start and finish a `Run` around a block |
| `init_run` | function | Start a `Run` you finish yourself |
| `experiment` | decorator | Wrap a function in a `Run` |
| `episode` | decorator | Wrap a function in an episode `Rollout` inside the active run |
| `log` | function | Log metrics to the active run |
| `score` | function | Post a score from the active run |
| `success` | function | Emit a success signal from the active run |
| `finish` | function | Finish the active run |
| `get_current_run` | function | Return the active `Run` or `None` |
| `get_current_rollout` | function | Return the active episode `Rollout` or `None` |
| `Run` | class | A run; `log`, `score`, `success`, `rollout`, `finish` |
| `Metric` | dataclass | A metric with full metadata for `log_metrics` |
| `GatewayRunError` | exception | Run creation or successful completion could not meet its durability contract |

<Info>
  `gatewaysdk.Rollout` is a **different** class from the episode returned by `run.rollout()`. The top-level `Rollout` (from `gatewaysdk.types`) is trainer rollout metadata with no scoring methods. The scoring episode documented on [Score from an episode](#score-from-an-episode-with-rolloutscore) is never imported directly: you always get it from `run.rollout(...)`.
</Info>

## Start a run with the `run` context manager

`run()` opens a run, yields a `Run`, and finishes it on exit, completing normally or marking `failed` when the body raises.

```python theme={null}
def run(name: Optional[str] = None, **kwargs) -> ContextManager[Run]
```

Every keyword after `name` is forwarded to the `Run` constructor.

```python theme={null}
import gatewaysdk

with gatewaysdk.run("gepa-v1", config={"lr": 0.01}, tags=["search"]) as r:
    for epoch in range(100):
        loss = train_one_epoch()
        r.log({"loss": loss}, step=epoch)
```

## Start a run manually with `init_run`

`init_run()` constructs a `Run`, starts it on the backend, and returns it. Call `finish()` yourself: no context manager wraps it.

```python theme={null}
def init_run(
    name: Optional[str] = None,
    *,
    project_id: Optional[str] = None,
    heartbeat_interval: int = 30,
    source: str = "sdk",
    labels: Optional[Dict[str, Any]] = None,
    config: Optional[Dict[str, Any]] = None,
    tags: Optional[List[str]] = None,
    **kwargs,
) -> Run
```

```python theme={null}
import gatewaysdk

r = gatewaysdk.init_run("manual-run", config={"seed": 7})
r.log({"step": 1})
r.finish(state="completed")
```

### Run constructor parameters

`run()`, `init_run()`, `experiment()`, and `Run(...)` all accept the same constructor keywords.

| Parameter | Type | Default | Description |
| - | - | - | - |
| `name` | `str` | `f"run-{uuid}"` | Human-readable run name |
| `project_id` | `str` | `$GATEWAY_PROJECT_ID` | Project the run belongs to. Required; raises `ValueError` when unset |
| `heartbeat_interval` | `int` | `30` | Seconds between heartbeats. A missed heartbeat marks the run `crashed` |
| `source` | `str` | `"sdk"` | Source identifier, for example `"gepa"` or `"trainer"` |
| `labels` | `dict` | `None` | Arbitrary key-value labels for filtering |
| `config` | `dict` | `None` | Hyperparameters recorded for reproducibility |
| `tags` | `list[str]` | `None` | Simple string tags for categorization |
| `api_url` | `str` | `$GATEWAY_HOST` | Backend URL. Falls back to `http://localhost:3000` |
| `api_key` | `str` | `$GATEWAY_SECRET_KEY` | Secret API key |

<Info>
  A missing project ID, a failed create request, or an invalid backend response raises `gatewaysdk.GatewayRunError`. The SDK does not fabricate a local run or run the body as if tracking had started. On create, it also records the current git SHA, branch, dirty state, and a config hash for reproducibility.
</Info>

## Decorate a function as an experiment or episode

`experiment()` wraps a whole function in a run; `episode()` wraps a function in an episode inside the already-active run.

```python theme={null}
def experiment(
    name: Optional[str] = None,
    *,
    inject_run: bool = True,
    **run_kwargs,
) -> Callable[[F], F]

def episode(
    name: Optional[str] = None,
    *,
    name_fn: Optional[Callable[..., str]] = None,
    inject_rollout: bool = False,
) -> Callable[[F], F]
```

| Parameter | Type | Default | Description |
| - | - | - | - |
| `name` | `str` | function name | Run or episode name |
| `inject_run` | `bool` | `True` | Pass the `Run` as the decorated function's first argument (`experiment` only) |
| `name_fn` | `callable` | `None` | Build the episode name from the call arguments (`episode` only) |
| `inject_rollout` | `bool` | `False` | Pass the `Rollout` as the decorated function's first argument (`episode` only) |
| `**run_kwargs` | n/a | n/a | Forwarded to the `Run` constructor (`experiment` only) |

Both decorators support synchronous and asynchronous functions. `episode()` raises `RuntimeError` when no run is active.

```python theme={null}
import gatewaysdk

@gatewaysdk.experiment("gepa-v1", config={"lr": 0.01})
def train(run):
    for candidate in population:
        evaluate(candidate)

@gatewaysdk.episode(name_fn=lambda c: f"candidate-{c.id}")
def evaluate(candidate):
    result = agent.execute(candidate.task)
    gatewaysdk.score("rubric", grade(result))

train()
```

## `Run` methods and properties

A `Run` tracks metrics, spawns episodes, posts scores, and finishes with a durable final state.

### Read-only properties

| Property | Type | Description |
| - | - | - |
| `id` | `str \| None` | Backend run ID, available after start |
| `run_number` | `int` | Auto-incrementing per-project run number |
| `name` | `str` | Run name |
| `state` | `str` | One of `pending`, `running`, `completed`, `failed`, `cancelled`, `crashed` |
| `project_id` | `str \| None` | Resolved project ID |

### Log metrics with `run.log()`

```python theme={null}
def log(
    self,
    metrics: Union[Dict[str, Any], str],
    value: Optional[float] = None,
    *,
    step: Optional[int] = None,
    labels: Optional[Dict[str, str]] = None,
    unit: str = "",
    metric_type: str = "gauge",
) -> None
```

| Parameter | Type | Default | Description |
| - | - | - | - |
| `metrics` | `dict` or `str` | required | A `{name: value}` dict, or a single metric name |
| `value` | `float` | `None` | Required only when `metrics` is a single name |
| `step` | `int` | `None` | Training step or iteration number |
| `labels` | `dict` | `None` | Per-metric labels for filtering |
| `unit` | `str` | `""` | Unit of measurement, for example `"ms"` or `"tokens"` |
| `metric_type` | `str` | `"gauge"` | Visualization hint: `"gauge"`, `"counter"`, or `"histogram"` |

Metrics are buffered and flushed roughly once per second, so they appear on the run page while the run is still going. Logging to a finished run logs a warning and no-ops.

```python theme={null}
with gatewaysdk.run("eval") as r:
    r.log({"accuracy": 0.92, "f1": 0.88}, step=10)
    r.log("latency_ms", 240.0, step=10, unit="ms")
```

### Other logging methods

| Method | Signature | Purpose |
| - | - | - |
| `log_metrics` | `log_metrics(metrics: List[Metric]) -> None` | Log a batch of `Metric` objects, each with full metadata |
| `flush` | `flush() -> None` | Force an immediate flush of buffered metrics |
| `log_state` | `log_state(state: Dict[str, Any]) -> None` | Stream rich structured state (candidates, trees, Pareto fronts) for live visualization |

`Metric` is a dataclass with fields `name: str`, `value: float`, `step: Optional[int] = None`, `unit: str = ""`, `metric_type: str = "gauge"`, and `labels: Dict[str, str] = {}`.

### Create an episode with `run.rollout()`

```python theme={null}
def rollout(self, name: Optional[str] = None) -> Rollout
```

`rollout()` returns an episode `Rollout` (a context manager) that carries its own session ID. Traces produced inside the episode are grouped under that session, and episode-level `score()` and `success()` calls target it. See [Score from an episode](#score-from-an-episode-with-rolloutscore).

### Score a session or trace with `run.score()`

```python theme={null}
def score(
    self,
    name: str,
    value: Union[float, str],
    *,
    trace_id: Optional[str] = None,
    session_id: Optional[str] = None,
    observation_id: Optional[str] = None,
    data_type: str = "NUMERIC",
    comment: Optional[str] = None,
    metadata: Optional[Dict[str, Any]] = None,
) -> None
```

| Parameter | Type | Default | Description |
| - | - | - | - |
| `name` | `str` | required | Score name, for example `"rubric"` or `"accuracy"` |
| `value` | `float` or `str` | required | Numeric for `NUMERIC`/`BOOLEAN`, string for `CATEGORICAL` |
| `trace_id` | `str` | `None` | Trace to target when no session is available; auto-detected from the active OpenTelemetry span |
| `session_id` | `str` | `None` | Session to target; auto-detected from the active episode |
| `observation_id` | `str` | `None` | Span ID for a trace-targeted score; ignored for session-targeted scores |
| `data_type` | `str` | `"NUMERIC"` | `"NUMERIC"`, `"CATEGORICAL"`, or `"BOOLEAN"` |
| `comment` | `str` | `None` | Optional explanation of the score |
| `metadata` | `dict` | `None` | Optional metadata merged with run context (`run_id`, `run_name`, `source`) |

Each score targets exactly one session or trace: the API enforces this XOR. When a session is available (an explicit `session_id` or an active episode), the session wins even if a trace is also present, and session-targeted scores need no trace. Only without a session does `score()` use `trace_id` or fall back to the current span; with neither target, the score is skipped with a warning.

Scores post in a background thread, and `finish()` waits for delivery. Scoring on a finished run warns and no-ops.

```python theme={null}
with gatewaysdk.run("eval") as r:
    with r.rollout("task-1") as episode:
        result = agent.execute(task)
        episode.score("rubric", 0.85)                 # targets the session
        r.score("accuracy", 0.95, trace_id="abc123")  # explicit trace target
```

### Feed a Success Metric with `run.success()`

`success()` is a thin wrapper over `score()` that tags the score as a success signal so the platform can roll it into an agent's Success Metric.

```python theme={null}
def success(
    self,
    name: str,
    value: Union[bool, float] = True,
    *,
    agent_id: Optional[str] = None,
    metric_name: Optional[str] = None,
    trace_id: Optional[str] = None,
    session_id: Optional[str] = None,
    observation_id: Optional[str] = None,
    data_type: Optional[str] = None,
    comment: Optional[str] = None,
    metadata: Optional[Dict[str, Any]] = None,
) -> None
```

| Parameter | Type | Default | Description |
| - | - | - | - |
| `name` | `str` | required | Signal name. Match it to the metric's configured SDK signal |
| `value` | `bool` or `float` | `True` | `True`/`False` for a pass/fail signal, or a number |
| `agent_id` | `str` | `None` | Agent the signal belongs to (`gateway.agent.id`). Defaults to the trace's agent when omitted |
| `metric_name` | `str` | `None` | Success Metric name; defaults to `name` |
| `trace_id`, `session_id`, `observation_id` | `str` | `None` | Linkage; auto-detected like `score()` |
| `data_type` | `str` | inferred | Overrides the inferred `"BOOLEAN"`/`"NUMERIC"` |
| `comment` | `str` | `None` | Optional comment |
| `metadata` | `dict` | `None` | Extra metadata merged into the signal |

A boolean `value` records a `1`/`0` `BOOLEAN` score; a number records a `NUMERIC` score. The call adds `gateway.success.signal = True` and `gateway.success.metricName` (and `gateway.agent.id` when `agent_id` is set) to the score metadata. Target resolution is identical to `score()`.

```python theme={null}
with gatewaysdk.run("support-eval") as r:
    with r.rollout("ticket-42") as episode:
        resolved = agent.resolve(ticket)
        episode.success(
            "resolved_ticket",
            resolved,
            agent_id="support-triage",
            metric_name="Resolution rate",
        )
```

### Finish a run with `run.finish()`

```python theme={null}
def finish(self, state: RunState = "completed") -> None
```

`state` is one of `"completed"`, `"failed"`, or `"cancelled"`. Finishing stops the heartbeat, drains buffered metrics, waits for in-flight score threads, and sends the final state.

<Info>
  `completed` is confirmed only after the SDK drains every accepted metric, joins accepted score requests, and the backend accepts the final state. If a metric or score was lost, or the final request is rejected or cannot be confirmed, `finish()` sets the run state to `failed` and raises `GatewayRunError`. Explicit `failed` and `cancelled` finalization stays best-effort so a telemetry failure never masks the original body exception.
</Info>

```python theme={null}
try:
    r.finish(state="completed")
except gatewaysdk.GatewayRunError as error:
    handle_delivery_failure(error)
```

### Propagate remote cancellation

`set_on_cancel(callback)` registers a zero-argument callback fired when a "Cancel Run" click in the dashboard reaches the run through its heartbeat. Trainers set it to an algorithm's cancel hook so a UI cancel stops the running algorithm.

```python theme={null}
def set_on_cancel(self, callback: Callable[[], None]) -> None
```

## Score from an episode with `rollout.score()`

An episode `Rollout` (returned by `run.rollout()`) captures its own session ID, and captures the first trace ID that starts inside its context. Its `score()` and `success()` methods auto-fill both, then delegate to the parent run.

```python theme={null}
def score(self, name: str, value: Union[float, str], **kwargs) -> None
def success(self, name: str, value: Union[bool, float] = True, **kwargs) -> None
```

Every keyword accepted by `Run.score()` and `Run.success()` passes through. The episode supplies `session_id` from itself and `trace_id` from its first span unless you override them.

| Property | Type | Description |
| - | - | - |
| `session_id` | `str` | Session ID injected into traces in this episode (`ep-{uuid}`) |
| `trace_id` | `str \| None` | Trace ID captured from the episode's first span |
| `name` | `str \| None` | Episode name |
| `run` | `Run` | Parent run |

```python theme={null}
with gatewaysdk.run("gepa-v1") as r:
    for candidate in population:
        with r.rollout(f"candidate-{candidate.id}") as episode:
            result = agent.execute(candidate.task)
            episode.score("rubric", grade(result))
```

## Module-level shortcuts post to the active run

Four module-level functions act on whichever run is currently active, so shared helper code can log or score without threading a `Run` through every call.

```python theme={null}
def log(metrics: Dict[str, Any], step: Optional[int] = None, **kwargs) -> None
def score(name: str, value: Union[float, str], **kwargs) -> None
def success(name: str, value: Union[bool, float] = True, **kwargs) -> None
def finish(state: RunState = "completed") -> None
```

| Function | No active run | Delegates to |
| - | - | - |
| `gatewaysdk.log` | raises `RuntimeError` | `Run.log` |
| `gatewaysdk.score` | raises `RuntimeError` | `Run.score` |
| `gatewaysdk.success` | raises `RuntimeError` | `Run.success` |
| `gatewaysdk.finish` | returns silently (no-op) | `Run.finish` |

<Info>
  `log`, `score`, and `success` raise `RuntimeError` when no run is active: call them inside `with gatewaysdk.run(...)`, under `@experiment`, or after `init_run()`. `finish` is the exception: it quietly does nothing when no run is active.
</Info>

```python theme={null}
import gatewaysdk

with gatewaysdk.run("eval"):
    gatewaysdk.log({"accuracy": 0.92}, step=1)
    gatewaysdk.success("passed", True, agent_id="triage")
```

### Reach the active run or episode

```python theme={null}
def get_current_run() -> Optional[Run]
def get_current_rollout() -> Optional[Rollout]
```

Both return `None` when nothing is active, which makes logging optional in shared code.

```python theme={null}
run = gatewaysdk.get_current_run()          # Run or None
episode = gatewaysdk.get_current_rollout()  # Rollout or None
```

## Map score values to platform concepts

Every `score()` and `success()` call creates a Gateway Score entity. The `data_type` decides how the platform reads the value.

| `data_type` | `value` | How the platform reads it |
| - | - | - |
| `NUMERIC` | a number | A numeric score, averaged or summed in aggregations |
| `BOOLEAN` | `1.0` / `0.0` | A pass/fail signal; `success(name, True)` records `1.0`, `False` records `0.0` |
| `CATEGORICAL` | a string | A labeled category; passing is decided by the score config |

`success()` infers `BOOLEAN` from a `bool` and `NUMERIC` from a number unless you set `data_type`. A boolean signal drives a **Pass rate** Success Metric; a numeric signal drives an **Average**, **Count**, or **Sum**.

The `agent_id` argument sets `gateway.agent.id` on the score. Per-agent cost on the Agents page and per-agent Success Metric roll-ups key off the same ID, so a `success()` call tagged with `agent_id` reports next to that agent's spend.

<Info>
  The signal `name` and the Success Metric name are separate. Set `metric_name` when they differ so an unrelated same-named score is never counted toward the metric.
</Info>

## Register and read A/B experiments

The experiments namespace assigns users to variants deterministically and logs exposures. Access it through the `gatewaysdk.experiments` proxy; no client instance is required.

### Initialize the backend connection

`init_experiments()` starts the background exposure logger and creates the global experiment client. Call it once at startup for backend-synced experiments; skip it to run local-only (in-memory, no exposure logging).

```python theme={null}
def init_experiments(
    api_endpoint=None,
    public_key=None,
    secret_key=None,
    batch_size=100,
    flush_interval=1.0,
    api_base_url=None,
) -> ExperimentClient
```

| Parameter | Type | Default | Description |
| - | - | - | - |
| `api_endpoint` | `str` | `None` | Full URL of the exposures batch API. Derived from `api_base_url` when omitted |
| `public_key` | `str` | `None` | Gateway public key (`pk-lf-...`) |
| `secret_key` | `str` | `None` | Gateway secret key (`sk-lf-...`) |
| `batch_size` | `int` | `100` | Maximum exposures per batch |
| `flush_interval` | `float` | `1.0` | Seconds between exposure flushes |
| `api_base_url` | `str` | `None` | Base URL for fetching and syncing experiments. Derived from `api_endpoint` when omitted |

### The `experiments` namespace methods

| Method | Signature | Purpose |
| - | - | - |
| `register` | `register(key, variants, traffic_percent=100, sync_to_backend=True)` | Define an experiment in code. Needs at least 2 variants; `traffic_percent` is 0–100. Syncs to the backend in a background thread |
| `get_variant` | `get_variant(experiment_key, user_id) -> Optional[Variant]` | Deterministically assign a user, set the trace context, and log an exposure. Returns `None` outside the traffic allocation or for an unregistered key |
| `is_registered` | `is_registered(experiment_key) -> bool` | Report whether an experiment is registered locally |
| `list_experiments` | `list_experiments() -> List[str]` | List registered experiment keys |
| `fetch` | `fetch(experiment_key, auto_register=True) -> Optional[ExperimentConfig]` | Synchronously fetch one experiment from the backend and register it locally |
| `fetch_all` | `fetch_all(auto_register=True) -> List[ExperimentConfig]` | Fetch every backend experiment |
| `configure_backend` | `configure_backend(api_base_url, public_key, secret_key)` | Point the client at a backend without `init_experiments()` |

`register()` takes each variant as a dict with `key`, `weight`, and `config`. `get_variant()` returns a `Variant` dataclass with `key: str`, `weight: int`, and `config: Dict[str, Any]`.

```python theme={null}
import os
import gatewaysdk

gatewaysdk.init_experiments(
    api_base_url=os.environ["GATEWAY_HOST"],
    public_key=os.environ["GATEWAY_PUBLIC_KEY"],
    secret_key=os.environ["GATEWAY_SECRET_KEY"],
)

gatewaysdk.experiments.register(
    key="prompt-test-v1",
    variants=[
        {"key": "control", "weight": 50, "config": {"prompt": "You are a helpful assistant."}},
        {"key": "treatment", "weight": 50, "config": {"prompt": "You are a concise assistant."}},
    ],
    traffic_percent=100,
)

variant = gatewaysdk.experiments.get_variant("prompt-test-v1", user_id="user-123")
if variant:
    system_prompt = variant.config["prompt"]
```

When the backend is configured, `get_variant()` uses server-side balanced assignment for exact proportional allocation and falls back to local hash-based assignment if the backend is unreachable. The assigned variant is written to the OpenTelemetry span context, so traces produced afterward carry the experiment and variant.

## Related pages

* [Runs & Experiments](/sdk/run-decorator): the narrative guide to the same API.
* [Evaluation & Replay](/evaluation): read the scores back against a pinned world version.
* [Tracing](/tracing): produce the spans that `score()` links to.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.