> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Run a task suite

> runSessions opens one warm container per task, runs your agent against each in parallel, grades them, and gates the mean reward with failUnder.

`runSessions` runs a regression suite. It opens one container per task, runs your agent against each of them in parallel, grades every session, closes them warm, and returns one report with a pass or fail verdict.

Each task needs its own container: two agents sharing one world state would corrupt both episodes.

## Run every scenario and gate the mean

```typescript theme={null}
import { runSessions } from "@withgateway/sdk/worlds";

const report = await runSessions({
  world: "acme-billing@main",
  concurrency: 4,
  failUnder: 0.7,
  agent: async ({ toolkit, instruction }) => {
    return myAgent(toolkit.impls, toolkit.schemas, instruction);
  },
});

console.log(report.mean.toFixed(3), report.passed);
for (const result of report.results) {
  console.log(result.task, result.reward, result.error ?? "");
}

process.exit(report.passed ? 0 : 1);
```

With no `tasks` given, every bundled task in the world runs. The world is resolved once for the whole run, so every session pins the same version and a push mid-suite cannot split the results across two worlds.

## What the handler receives

Your `agent` function is called once per task with a session that is already ready and a toolkit already bound to it. It runs inside `session.withTraceContext`, so the spans it records after `tracing.init()` group under that session's id.

| Field | Type | What it is |
| - | - | - |
| `session` | `WorldSession` | The live session, if you need `state()` or `seed()`. `runSessions` opens the tools surface. When the agent speaks HTTP, open the session yourself with `surfaces: ["api"]` |
| `toolkit` | `WorldToolkit` | `schemas`, `impls` and `specs` for this world's tools |
| `instruction` | `string` | What this task asks for |
| `task` | `string` | The task's label, matching `SessionRunResult.task` |
| `taskRef` | `WorldTaskRef` | The task as you gave it: a name, a stored ref, or the inline spec |

Grading, closing and aggregation are handled for you. Return whatever you like from the handler; the reward comes from the task's grader, not from the return value.

## Every option

| Option | Type | Default | Meaning |
| - | - | - | - |
| `world` | `string` | required | Slug or `slug@ref` |
| `agent` | `(ctx) => Promise<unknown>` | required | Your agent, called once per task |
| `tasks` | `WorldTaskRef[]` | every bundled task | Names, stored refs and inline specs, mixed freely |
| `concurrency` | `number` | `3` | Sessions open at once, minimum 1 |
| `failUnder` | `number` | none | The bar `mean` must meet for `passed` to be true |
| `versionId` | `string` | the world's head | Pin every session to one version |
| `keepWarm` | `boolean` | `true` | Hand containers back warm for the next run |
| `onResult` | `(result) => void` | none | Called as each task finishes, for live output |
| `options` | `WorldsClientOptions` | environment | Host and keys |

## The report

`runSessions` resolves to a `SessionRunReport`.

| Field | Type | What it is |
| - | - | - |
| `results` | `SessionRunResult[]` | One entry per task, sorted by task name |
| `mean` | `number` | Mean reward over every task, counting failures as zero |
| `passed` | `boolean` | `mean >= failUnder`, or `true` when no gate was set |

Each `SessionRunResult` carries `task`, `sessionId`, `reward`, `rewards`, and an `error` string when the task threw instead of grading.

<Info>
  A task that throws scores zero and records why. The other tasks keep their results.
</Info>

## Mix bundled, stored and inline tasks

`tasks` accepts all three kinds of reference in the same array.

```typescript theme={null}
const report = await runSessions({
  world: "acme-billing@main",
  tasks: [
    "refund-double-charge",                    // a bundled scenario
    { id: "wt_4f1c", version: 2 },             // a stored task, pinned
    {                                          // an inline task
      name: "pay-inv-7",
      instruction: "Mark invoice INV-7 as paid.",
      seed: { rows: { invoices: [{ id: "INV-7", status: "open" }] } },
      grader: {
        kind: "assertions",
        checks: [{ entity: "invoices", where: { id: "INV-7", status: "paid" } }],
      },
    },
  ],
  concurrency: 3,
  agent: async ({ toolkit, instruction }) => myAgent(toolkit.impls, instruction),
});
```

An unnamed inline task is labelled `inline:<hash>` in the results.

## Stream results as they land

`onResult` fires as each task finishes.

```typescript theme={null}
const report = await runSessions({
  world: "acme-billing@main",
  failUnder: 0.8,
  onResult: (r) =>
    console.log(`${r.reward === null ? "  ?  " : r.reward.toFixed(2)}  ${r.task}${r.error ? `  ${r.error}` : ""}`),
  agent: async ({ toolkit, instruction }) => myAgent(toolkit.impls, instruction),
});
```

<Info>
  `runSessions` runs **your** agent against real containers. To run the platform's own agent instead, dispatch a run config with [`world.dispatch`](/sdk-ts/worlds#dispatch-a-hosted-run), or gate a pull request with [`validateWorld`](/sdk-ts/ci).
</Info>

## Where to go next

* [Gate a pull request](/sdk-ts/ci) when the world source itself is what changed.
* [World sessions](/sdk-ts/sessions) for the methods on the `session` your handler receives.
* [Run many sessions at once](/worlds/parallel) for the same idea from the terminal and from Python.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.