> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluation (LLM-as-Judge)

> How a world's runs are scored - write the criteria a judge grades against, point a check at the runs you care about, and read the scores it produces.

Evaluation scores runs. You write the criteria, point a check at the runs you care about, and Surface Area runs a judge over each match and attaches a score. The dashboard calls the feature **Checks**; the scores appear under the **Results** tab.

One check grades both a [rollout](/glossary#rollout) inside a [world](/worlds) and a production session, so scores from the two are comparable.

<img src="https://mintcdn.com/surface-d3d890e1/I9MKHQA4beHtjYQ2/screenshots/evals-home.png?fit=max&auto=format&n=I9MKHQA4beHtjYQ2&q=85&s=2af9ee068555a132c7a0523412b1b632" alt="The Evals home with checks and results" width="2880" height="1800" data-path="screenshots/evals-home.png" />

*The Evals surface: checks on the left, the scores they produce under Results.*

## How evaluation works

A check watches for runs that match a filter, sends each one to a judge with your criteria, and writes the verdict back as a score on the run.

The default judge is an LLM. You give it a prompt describing what "good" looks like, and it returns a score plus its reasoning. Deterministic checks score with code instead of a model.

Each score has a name, a value, and a data type, and attaches to the trace or session that was graded. Scores appear as columns in [Tracing](/tracing), roll up into a run's pass rate, and are what a release gate reads.

<Info>
  LLM and agent checks need a default evaluation model. Set one with **Default model** on the Checks page before you create a check -- a check with no model fails fast. The model needs a provider connection (for example an Anthropic or OpenAI key under **Settings → LLM connections**).
</Info>

## Create a check

Open **Checks** from the project navigation and select **New check**.

<Steps>
  <Step title="Name the check and pick how it runs">
    Give the check a score name -- the name its scores carry. Then choose an execution mode:

    | Mode | What it does |
    | - | - |
    | LLM Judge | An LLM reads your criteria and returns a score with reasoning. The default. |
    | Agent Judge | An agent with tools investigates the run before scoring, for checks that need more than a single prompt. |
    | Deterministic | A built-in rule scores the run with no model -- for example exact match, contains, regex, or valid JSON. |
    | Custom Code | Your own Python scores the run, with its own inputs per check. |
  </Step>

  <Step title="Write the criteria">
    For an LLM or agent judge, write the criteria the judge grades against. Describe a passing run in plain language. The judge reads your criteria together with the run's input and output.

    Deterministic checks take a rule type instead, such as **Exact Match**, **Contains**, **Regex Match**, or **Valid JSON**. Custom code checks take a Python script.
  </Step>

  <Step title="Choose what the check runs on">
    Pick the **target**: **trace** to grade individual runs, or **session** to grade a whole conversation or workflow. The target decides which data the check can see.

    Add a **filter** to narrow the target down -- by the trace's `environment` field, a tag, metadata, or any other condition. Leave the filter empty to grade everything. A preview shows the most recent matching runs while you build the filter.
  </Step>

  <Step title="Set the time scope">
    The **time scope** decides which runs the check grades:

    | Scope | Behavior |
    | - | - |
    | New | Grades every future run that matches the filter. |
    | Existing | Grades the runs already in your project that match the filter. |

    <Info>
      Running a check on **Existing** runs is a one-time backfill. Pick a tight filter before you grade existing data so you do not backfill more runs than you mean to.
    </Info>
  </Step>

  <Step title="Map your data into the judge">
    Variable mapping connects the placeholders in your prompt to fields on the run. A prompt that uses `{{input}}` and `{{output}}` needs each variable pointed at a column -- for example the trace input and the trace output.

    Surface Area infers a default mapping from your prompt's variables. Adjust it when the field you want to grade is not the top-level input or output.
  </Step>
</Steps>

<Info>
  The Checks page has a **Build checks with the agent** card. Tell the agent the behavior to verify, and it drafts the check, runs it, and keeps the results on the page.
</Info>

## Choose a score's data type

Each check defines a **score config** -- the shape of the value it writes. Reuse an existing config or create a new one.

| Data type | What it captures |
| - | - |
| Numeric | A number in a range you set, for example 0 to 1 |
| Categorical | One label from a set you define |
| Boolean | A pass/fail (true/false) verdict |

A score config also carries an optional **description** shown to human annotators, so machine scores and human review use the same definition. Sharing one config across checks keeps the scoring schema consistent.

<Info>
  A categorical judge can correctly return no label when nothing applies. Surface Area treats a missing score as its own state, never as zero -- absent and zero stay distinct in tables and charts.
</Info>

## Produce several scores from one check

One check can write several scores from a single judge call. Add multiple scores to the check, each with its own name, criteria, and score config.

The judge reads the run once and returns every score together. Use multi-score checks when one run needs grading on several independent dimensions -- accuracy, relevance, and tone, for example.

## Tune when and how often a check runs

**Sampling** sets the fraction of matching runs to grade, from 0 to 1. Lower it to spend less on high-volume traffic; 1 grades every match.

**Delay** holds a new run for a short window before grading, so late-arriving spans are included. A **session** check waits a configurable amount of time after the last activity before it grades the session as complete.

**Model override** gives one check a different provider, model, or model parameters than the project default.

<Info>
  The model a check uses drives its cost. A check left on the default evaluation model at full sampling on high-volume traffic can run a large number of judge calls. Start new checks at a lower sampling rate and raise it once you trust the check.
</Info>

## Stack checks into multi-tier evaluations

A multi-tier evaluation runs an expensive check only on the runs a cheaper check flags first. Stack checks you have already built into tiers, put a gate between them, and a run escalates to the next tier only when the gate's condition holds.

Open the **Multi-tier** view on the Checks page and select **Build a multi-tier eval** to stack existing checks onto the canvas.

<img src="https://mintcdn.com/surface-d3d890e1/I9MKHQA4beHtjYQ2/screenshots/multi-tier-builder.png?fit=max&auto=format&n=I9MKHQA4beHtjYQ2&q=85&s=ae8b3a002c94d453443dda0357f36634" alt="The multi-tier eval builder" width="2880" height="1800" data-path="screenshots/multi-tier-builder.png" />

*The multi-tier builder: stack checks into tiers and join them with gates.* The first tier is labeled **Runs first**; every tier you add below it is a **Then run** tier reached through a gate.

The first tier grades every matching run, exactly like a standalone check. Each later tier reads the scores from the tier above and runs only when its gate matches, so the expensive judge never touches a run the cheap check already cleared.

A gate compares the previous tier's scores:

| Gate mode | What it checks |
| - | - |
| Score | One comparison against a single score, such as `quality < 0.5`. |
| Multiple | Several comparisons joined with AND or OR. |
| Code | A condition you describe in plain language -- Surface Area writes the rule for logic a plain comparison can't express, such as "when one score is lower than another." |

The checks in a tier share one target -- all traces, or all sessions. Later tiers ignore their own filter and sampling settings; the gate is their only trigger. Choose per gate whether a missing score skips the next tier or runs it anyway.

The Assistant can wire multi-tier checks for you -- tell it "only run the detailed check when the quick one fails," and it links the tiers and gates. See [the Assistant](/worlds/assistant).

<Info>
  Put your cheapest, fastest check first -- a deterministic rule or a small model -- and gate the expensive LLM or agent judge behind it. Later tiers run only on gate matches and skip their own sampling, so the expensive judge grades a fraction of your traffic.
</Info>

## Read the scores a check produces

The **Results** tab on the Checks page lists the scores the project's checks have written. In [Tracing](/tracing), each evaluator's scores get their own column on the sessions and traces tables, so you can scan quality across many runs and filter by score value.

Open a scored run to see the score next to the run's input and output -- exactly what the judge read.

## Keep checks healthy

The Checks table shows each check's health and recent results, and gives you the controls to manage it.

**Activate / Deactivate** turns a check on or off from its row. A deactivated check stops grading new runs but keeps its history and scores.

**Rerun failed** retries only the runs whose grading errored; **Rerun all** re-grades every matching run. Both are in the row's menu.

<Info>
  Editing and saving a check triggers a fresh grading batch. To grade runs whose grading was cancelled, save the check or rerun it.
</Info>

## Benchmarks versus checks

A **check** scores runs that already happened: a judge reads each matching trace or session and writes a score. A **benchmark** produces the runs: a world with tasks, each task a prompt plus SQL verifiers over the world's end state, run k times per model.

| | Check (this page) | Benchmark |
| - | - | - |
| Input | Existing traces and sessions, from a world or production | A world's tasks, run on demand |
| Grader | An LLM judge, or code, over the trace | SQL verifiers over the world's state; a rubric on a stored task |
| Output | A score on each trace or session | Per-task mean reward, pass rate, samples, null scores |
| Where | Checks and Results tabs | `gateway benchmarks run` / `results`; the world page's Results tab |

A benchmark's rollouts are sessions, so a check can score them too. Build one with [gateway benchmarks](/cli/benchmarks); the task contract is on [Benchmarks on worlds](/worlds/benchmarks). A [harbor benchmark](./benchmarks-harbor) grades a container instead, runs on the benchmark runner, and reads back through the same results commands.

## Where evaluation connects to the rest of Surface Area

* [Worlds](/worlds) are where a candidate version is run. A check scores those runs the same way it scores production traffic. [Benchmarks on worlds](/worlds/benchmarks) turn a world's tasks into repeatable runs per model.
* [Scenarios & task sets](./data-task-sets) hold the situations a world runs, each with the conditions that decide whether it passed.
* [Agent Replay](./agent-replay) re-runs real sessions against a new version, and an evaluator grades a replay session the same way it grades a live one.
* [Tracing](/tracing) holds the sessions and traces every check grades, and shows the scores as columns you can filter and read.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.