> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Evals in CI/CD

> A GitHub Actions workflow that opens a session per scenario on the platform, runs the agent under test against each session's URL, grades the end state, and fails the pull request below a threshold.

**Every pull request gets a number.** The workflow below opens one live session per scenario, points the agent under test at each session's URL, grades what the world holds at the end, and fails the check when the mean reward is below your bar.

The agent talks to the session instead of the real vendor, so the platform sees every call it makes and grades the state it leaves behind. Nothing about the agent changes except a base URL and a header.

<Info>
  Every recipe here is plain workflow YAML around
  `npm install -g @withgateway/sdk` and the `gateway` command, the same shape
  [Run API worlds in CI](/worlds/ci) uses.
</Info>

## What the workflow needs

The workflow needs three repository secrets. The keys decide which project the worlds are read from, so a key pointed at a staging project keeps pull requests away from production data.

| Secret | What it is |
| - | - |
| `GATEWAY_HOST` | Your Surface Area URL, for example `https://withgateway.ai` |
| `GATEWAY_PUBLIC_KEY` | Project public key, `pk-lf-…` |
| `GATEWAY_SECRET_KEY` | Project secret key, `sk-lf-…` |

The CLI reads all three from the environment, so no `gateway auth login` step is needed. `gateway auth status` early in the job prints which host and key are in effect, so a missing secret shows up there instead of as a 401 later.

## The workflow

```yaml filename=".github/workflows/agent-evals.yml" theme={null}
name: agent-evals

on:
  pull_request:
  workflow_dispatch:

permissions:
  contents: read

env:
  GATEWAY_HOST: ${{ secrets.GATEWAY_HOST }}
  GATEWAY_PUBLIC_KEY: ${{ secrets.GATEWAY_PUBLIC_KEY }}
  GATEWAY_SECRET_KEY: ${{ secrets.GATEWAY_SECRET_KEY }}

jobs:
  agent-scenarios:
    name: agent · every scenario
    runs-on: ubuntu-latest
    timeout-minutes: 45
    env:
      WORLD: acme-ledger
      MIN_SCORE: "0.7"
    steps:
      - uses: actions/checkout@v4

      - uses: actions/setup-node@v4
        with:
          node-version: 22

      - name: Install the agent's own dependencies
        run: npm ci

      - name: Install the gateway CLI
        run: npm install -g @withgateway/sdk

      - name: Show which credentials are in effect
        run: gateway auth status

      - name: One session per scenario, agent against each URL, graded
        run: |
          gateway worlds run "$WORLD" \
            --agent ./agent/run.mjs \
            --all \
            --surface api \
            --timeout-min 20 \
            > eval-report.json || true

      - name: Fail below the threshold
        run: |
          jq -r '.results[] | "\(.task)\t\(.reward)\t\(.error // "")"' eval-report.json
          jq -r '"mean \(.mean)"' eval-report.json
          jq -e --argjson bar "$MIN_SCORE" '(.mean // 0) >= $bar' eval-report.json

      - uses: actions/upload-artifact@v4
        if: always()
        with:
          name: eval-report
          path: eval-report.json
          if-no-files-found: ignore

  world-tests:
    name: tests · the world itself
    runs-on: ubuntu-latest
    timeout-minutes: 20
    steps:
      - uses: actions/checkout@v4

      - uses: actions/setup-node@v4
        with:
          node-version: 22

      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"

      - run: pip install pytest

      - run: npm install -g @withgateway/sdk

      - name: Run the world's own test suite against this checkout
        run: |
          gateway worlds test ./worlds/acme-ledger \
            --out world-tests.json

      - uses: actions/upload-artifact@v4
        if: always()
        with:
          name: world-tests
          path: world-tests.json
          if-no-files-found: ignore
```

## The agent job runs the scenarios

`gateway worlds run <world> --all` asks the world for its task list, opens one live session per task, calls your agent with that session, and runs the task's own grader against the end state. Every session in the run is pinned to the same world version, resolved once, so a push mid-suite cannot split the results across two worlds.

`--surface api` makes the session answer over HTTP. Without it the session serves the tool surface only, and `session.api()` has no URL to return.

Your agent module default-exports `async (session, task) => void` and reads two values off the session:

```javascript filename="agent/run.mjs" theme={null}
export default async function agent(session, task) {
  const url = await session.api();          // this session's own base URL
  const headers = await session.apiHeaders(); // { Authorization: "Bearer …" }

  const open = await fetch(`${url}/v1/charges?status=disputed`, { headers });
  const { data } = await open.json();

  for (const charge of data) {
    await fetch(`${url}/v1/refunds`, {
      method: "POST",
      headers: { ...headers, "Content-Type": "application/json" },
      body: JSON.stringify({ charge: charge.id }),
    });
  }
}
```

A `.ts` agent works when `tsx` or `ts-node` is installed next to it; otherwise build to JavaScript and point `--agent` at the output.

The summary prints to stdout as JSON: `world`, `agent`, `mean`, and one `results` entry per scenario carrying `task`, `sessionId`, `reward`, `rewards` and an `error` string when the scenario threw instead of grading. Progress lines go to stderr, so redirecting stdout to a file keeps the report clean.

<Info>
  `gateway worlds run` exits 1 when **any** scenario threw, whatever the mean
  was. The `|| true` in the step above hands that decision to the threshold
  step instead, so one crashed scenario shows up as a zero in the mean rather
  than as a job that died before printing anything. Drop the `|| true` when you
  want a single crash to fail the pull request on its own.
</Info>

## The tests job checks the world

`gateway worlds test <path>` runs the suite the world ships for itself: the `tests/*.json` HTTP cases and the `tests/test_*.py` files described in [Test a world](/worlds/tests). It exits 1 when any case fails.

Pointing it at a path rather than a slug tests the tree in the pull request's checkout, which is the version that has not been published yet. With a `python3` of 3.12 or newer on the runner, the CLI serves the world from its bundled runtime and runs everything locally — no push, no version, no session.

**Install `pytest` alongside it.** Without it the Python half of the suite reports a skip, and a skip is not a pass.

<Info>
  The argument is read as a path first. A bare slug that matches a directory in
  the working directory tests that directory instead of the platform's copy, so
  keep the two distinct or pass `./worlds/<dir>` deliberately.
</Info>

## When the pull request changes the world itself

Neither job above is `gateway worlds ci`. `worlds test` runs the world's own assertions; `worlds ci` pushes the checked-out world, runs every scenario against the version it just pushed, and gates on the mean:

```bash theme={null}
gateway worlds ci ./worlds/acme-ledger --agent ./agent/run.mjs --min-score 0.7
```

Use it when the pull request edits the world's contract, handlers or data, so the runs grade the change being reviewed. It writes a summary table into `$GITHUB_STEP_SUMMARY`, takes `--summary-file` for the JSON, and defaults its branch to `GITHUB_HEAD_REF`, so each pull request's runs are tagged with its own branch.

[Run API worlds in CI](/worlds/ci) has the longer form of the same job — opening the session, masking the token, and polling until it is ready — plus the merge workflow that publishes on push.

## The knobs

### Parallelism

`gateway worlds run` opens its sessions one after another. There are three ways to widen it, in increasing order of effort:

| Approach | What it gives you | Cost |
| - | - | - |
| A GitHub Actions matrix, one scenario per entry | True parallelism across runners, each failure named in its own job | A runner and a CLI install per scenario |
| `runSessions` from the TypeScript SDK | `concurrency` sessions at once in one process, default 3 | Your agent gets a bound toolkit, not an HTTP URL |
| Your own pool over `openSession` | Both parallelism and `surfaces: ["api"]` | You handle grading, closing and aggregation |

An agent that needs a base URL calls `openSession(world, { task, surfaces: ["api"] })` in a pool of its own. [Run many sessions at once](/worlds/parallel) has both shapes, and [Run a task suite](/sdk-ts/run-sessions) covers `runSessions` in full.

### Warm sessions

`worlds run` closes each session as it finishes. A session closed warm goes back to a pool a later open can reuse: `gateway worlds session close <id> --keep-warm` from a shell, `reuseWarm: true` on `POST /api/public/world-sessions`, and the default for `runSessions`. Reuse is per version and per task.

### The threshold

Apply the score gate for `worlds run` in your workflow, in one of three places.

| Where | How |
| - | - |
| The workflow step | `jq -e --argjson bar "$MIN_SCORE" '(.mean // 0) >= $bar'` over the summary |
| `gateway worlds ci` | `--min-score 0.7`, which exits 1 under the bar |
| The TypeScript SDK | `runSessions({ failUnder: 0.7 })` sets `report.passed`; you still exit on it |

Pick a bar from a baseline run of the agent already on your main branch, not from a guess. A world's scenarios vary in difficulty, so the mean is only meaningful against the same world version.

### Where the results show up

**In the job.** The `eval-report.json` artifact holds every scenario's reward and error. Read it when the gate fails.

**In the dashboard.** Open **Worlds** and select the world. Its **Results**, **Runs** and **Evals** tabs carry what ran against it, and **Overview** lists live sessions while the job is still running. The section's **All results** tab is the matrix across every world in the project. The **Tests** tab holds the runs of the world's own suite, including the ones the platform ran on push.

## Where to go next

* [Run API worlds in CI](/worlds/ci) for the hand-rolled session job and the publish-on-merge workflow.
* [Test a world](/worlds/tests) for the case format and what the report says.
* [Run many sessions at once](/worlds/parallel) for the parallel shapes.
* [Simulations from real data](/use-cases/sims-from-real-data) for the other recipe.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.