> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Simulations for your users, from their real data

> The funnel from real traffic to a graded per-user run - get the rows out, redact them on the way in, give each tenant a world or each user a task, seed the session, run, grade, and look a user up again afterwards.

**Hand each of your customers a sealed copy of their own systems to run agents against.** The data comes from traffic you already have, redaction on the way in keeps personal fields out, and every run is graded against the state the agent left behind.

Six stages, each one a command.

| Stage | What happens | Command |
| - | - | - |
| 1. Get the rows out | Real traffic becomes rows: pull the raw tool calls, transform them with your code, ingest | `worlds data calls pull`, `worlds data ingest --transform`, `worlds session export`, `worlds data extract` |
| 2. Redact on the way in (optional) | Personal fields are hashed, dropped or shaped — off unless you ask | `--redact apply` |
| 3. Pick a shape | A world per tenant, or one world with per-user tasks | `worlds create`, `--overlay`, `--batch` |
| 4. Seed the run | One user's rows land in one live session | `worlds session seed`, a task's `seed` block |
| 5. Run and grade | The user's agent works, the grader reads the end state | `worlds session grade`, `worlds run` |
| 6. Look a user up | Support finds a hashed record from the plaintext they know | `worlds data query --resolve` |

<Info>
  Decide the redaction policy before the first import, not after. Rows already
  in a world were stored as given unless you asked otherwise, and re-hashing a
  digest that is already hashed breaks every join that depended on it.
</Info>

## Stage 1: get the real rows out

Three sources.

### Your agent's own runs, over OTLP

Instrument the agent with the `gatewaysdk` Python package and every LLM call, tool call and nested step lands in your project as a trace, grouped into sessions. Traces show **which situations to build**: the requests real users made, and what the vendor answered.

```python theme={null}
import gatewaysdk.tracing as tracing

tracing.init(service_name="billing-agent")
# ... your agent runs ...
tracing.shutdown()
```

Every tool call in those traces is a TOOL observation — the tool's name, its input and its output. A tool's output is not the entity's shape, so the route is pull, transform, ingest:

```bash theme={null}
gateway worlds data calls pull calls.jsonl --from traces --session <sessionId>     # 1. the raw calls, untouched; prints records per tool and each result shape
gateway worlds data ingest acme-crm calls.jsonl --transform shape.py --dry-run       # 2. your code shapes them; nothing written
gateway worlds data ingest acme-crm calls.jsonl --transform shape.py                 #    then for real
gateway worlds data query acme-crm accounts --limit 5                                # 3. prove it
```

`calls pull` takes `--trace` (repeatable), `--session` or `--since`, and writes `{tool, args, result, at, source}` records. The transform is your `transform(records) -> {entity: [rows]}` in Python or TypeScript; it runs in a subprocess, and the rows then go through exactly what `data import` does. `--map ingest.toml` is the no-code alternative, a declared mapping in which nothing is inferred. See [Pull, transform, ingest](/worlds/data#pull-transform-ingest).

To intercept the calls at the source instead, `captureToolCalls` from [the TypeScript SDK](/sdk-ts/ingest) (`capture_tool_calls` in Python) wraps the agent's tool functions so every real call against the vendor is written as the same record, with no change to the agent; `worlds data ingest` then reads the file. Traces remain the place to read *which situations* to build — the dashboard and [the REST API](/rest-api/resources) hold them, and [Export to Surface Area](/tracing/export) covers the endpoint, batching and flushing.

### A session's own call log

A session that already ran records every call it served. `worlds session export` writes the log as JSON Lines, oldest first.

```bash theme={null}
gateway worlds session export <sessionId> --out calls.jsonl
```

Each line carries `{args, completedAt, error, kind, result, seq, tool}`, where `kind` is `call`, `grade`, `seed`, `reset` or `close`. Reads carry rows in their `result`; writes carry the row the world should hold. `worlds data ingest` reads the file as it is, through the same transform or mapping.

### A vendor export, shaped by `data extract`

`worlds data extract` reads a file and writes rows already shaped for `data import`, against the world's own contract.

```bash theme={null}
gateway worlds data extract acme-crm calls.jsonl \
  --shape session-export --entity accounts --out accounts.jsonl
```

| `--shape` | What the input is |
| - | - |
| `session-export` | A `calls.jsonl` from `worlds session export`; rows come from the `result` of successful `call` records |
| `vendor-envelope` | The vendor's own responses, one per line or one document; rows sit under the envelope key the world declares |
| `tool-results` | JSON Lines carrying a `raw_output` object; rows come from its `results` (or `items`) list |

Extract drops null and empty fields, strips fields the entity does not declare and counts them, dedupes by primary key (`--merge fill` keeps the first row, `--merge upsert` the last), and **runs the world's `[redact]` policy over every row unless you pass `--no-redact`**. Its stderr summary is `{"read": N, "rows": M, "unknownFields": {…}, "redacted": {"dropped": n, "hashed": n}}`.

## Stage 2: redact on the way in (optional)

Redaction is optional. Every write path stores rows exactly as given unless you ask
otherwise, so a world seeded from already-sanitized data, or from data that carries nothing
personal, skips this stage. When the rows do carry personal fields, this stage hashes, drops
or shapes them.

The world's `connector.toml` declares one `[redact]` block, and every write path reads the same one.

```toml filename="connector.toml" theme={null}
[redact]
hash     = ["email", "full_name"]
drop     = ["ssn"]
preserve = ["account_number", "phone"]
round    = { latitude = 2, longitude = 2 }
```

`hash` stores a salted `sha256:<64 hex>` digest, `drop` removes the key entirely, `preserve` stores a format-preserving digest so an account number still looks like one, and `round` keeps a number and drops only its precision. [Keep real data out](/worlds/redaction) explains what each mode costs.

Every write path takes `--redact`, and **every one of them defaults to `off`**.

```bash theme={null}
# Rows from real traffic, redacted here rather than by hand.
gateway worlds data import acme-crm accounts.jsonl --redact apply
gateway worlds data ingest acme-crm calls.jsonl --transform shape.py --redact apply

# Check an already-redacted export before importing it.
gateway worlds data check acme-crm accounts.jsonl --redact refuse

# One live session only.
gateway worlds session seed <sessionId> accounts.jsonl --redact apply
```

`apply` runs the policy before the contract sees the row. `refuse` rejects a row that still carries plaintext in a redacted field, naming the entity, index and field. Both need the world to have a `connector.toml`; a world without one is refused up front rather than quietly storing plaintext.

<Info>
  Use `apply` for a world with `preserve` fields: a shaped digest looks like a
  real identifier by design.
</Info>

## Stage 3: pick the shape that matches what differs

| What differs between your customers | Shape | How it is built |
| - | - | - |
| The data and the configuration | **A world per tenant** | `gateway worlds create` once per customer |
| A few files on top of a shared base | **A world per tenant, from an overlay** | `gateway worlds create --overlay ./tenants/acme` |
| Only the goal | **One world, a task per end user** | One world, an inline or stored task spec |

Start with one world and per-user tasks. Move to a world per tenant only when customers need different data or a different contract: every world is a version history you then have to keep.

### One world per tenant

```bash theme={null}
gateway worlds create acme-crm \
  --from salesforce \
  --name "Acme CRM" \
  --rows acme-accounts.jsonl --entity accounts \
  -m "provisioned for Acme"
```

`--from` takes a shipped template slug, `workspace:<slug>` for a connector your organization published, or a world directory. `--overlay ./tenants/acme` lays a directory of files over the source before the tree is compiled, which is how one base contract becomes each tenant's variant.

Each tenant's world is stored as its schema tree, so one session serves both the world's tools and — when the connector declares routes — the vendor's HTTP paths.

Slugs are unique within a project, so prefix them with the tenant and settle the naming scheme before the first customer: a collision answers 409 rather than overwriting anything.

### Many tenants in one call

`--batch` takes a JSON Lines file (or `-` for standard input) of one world per line and prints one result line for each.

```jsonl filename="tenants.jsonl" theme={null}
{"slug": "acme-crm", "from": "salesforce", "name": "Acme CRM"}
{"slug": "globex-crm", "from": "salesforce", "name": "Globex CRM", "overlay": "./tenants/globex"}
{"slug": "initech-crm", "from": "workspace:our-crm", "name": "Initech CRM"}
```

```bash theme={null}
gateway worlds create --batch tenants.jsonl
```

A line may also carry `data` for the seed rows, as `{"rows": {"<entity>": [...]}}`. The command exits 0 when every world was created and 2 when any line failed, so a partial batch is visible in the exit code rather than only in the output.

Provisioning from your own backend instead is `POST /api/public/worlds`, the same request — see [Simulations for your users](/worlds/for-your-users#or-call-the-api-directly).

### One world, a task per end user

A task is an instruction, an optional seed and a grader. Supplying one at session open lets a single published version serve any number of end users, each with their own goal and starting rows.

```json filename="user-42.json" theme={null}
{
  "name": "mark-paid-INV-7",
  "instruction": "Mark invoice INV-7 as paid.",
  "seed": {
    "mode": "append",
    "redact": "apply",
    "rows": { "invoices": [{ "id": "INV-7", "status": "open" }] }
  },
  "grader": {
    "kind": "assertions",
    "checks": [{ "entity": "invoices", "where": { "id": "INV-7", "status": "paid" } }]
  },
  "metadata": { "externalUserId": "u_42" }
}
```

```bash theme={null}
gateway worlds session open acme-billing --task-file user-42.json --surface api
```

The `seed` block carries its own `redact` mode, so every session the task opens seeds the same way. `metadata` is echoed back on the session; put your own user id there.

Keep a task your users return to as a project resource rather than a request body:

```bash theme={null}
gateway worlds task create user-42.json --world acme-billing -n mark-paid-INV-7
gateway worlds task validate <taskId>
gateway worlds session open acme-billing --task-id <taskId> --surface api
```

`task validate` checks the seed and grader entities against the world's contract without opening a session. It catches a grader asserting on an entity the world does not define.

## Stage 4 and 5: seed, run, grade

<Steps>
  <Step title="Open a session for the run">
    ```bash theme={null}
    gateway worlds session open acme-crm --task-file run-8817.json --surface api
    ```

    The descriptor carries `sessionId`, `ready`, the surfaces it opened, and — with `--surface api` — an `api.url` and `api.token`.
  </Step>

  <Step title="Add this user's rows, if the task did not">
    ```bash theme={null}
    gateway worlds session seed <sessionId> user-8817-rows.jsonl --redact apply
    ```

    Seeding changes that one live copy and no version. `--mode replace` swaps the entities the file names; the default `--mode append` upserts by primary key.
  </Step>

  <Step title="Hand the user's agent the URL and the token">
    Point it at `api.url` with `Authorization: Bearer <api.token>`, as it would be pointed at the real vendor. Nothing else in its code changes.
  </Step>

  <Step title="Grade the end state">
    ```bash theme={null}
    gateway worlds session grade <sessionId>
    ```

    The answer is `{reward, rewards, raw}` — the headline number, the per-check breakdown, and whatever the grader returned.
  </Step>

  <Step title="Keep the ledger, then close">
    ```bash theme={null}
    gateway worlds session export <sessionId> --out run-8817.jsonl
    gateway worlds session close <sessionId>
    ```

    A session holds a running container and closes itself after thirty idle minutes. Closing it yourself returns the container immediately.
  </Step>
</Steps>

To run many users at once, `runSessions` from the TypeScript SDK opens `concurrency` sessions at a time and mixes bundled names, stored task refs and inline specs freely, so one call can run a task per user:

```typescript theme={null}
import { runSessions } from "@withgateway/sdk/worlds";

const report = await runSessions({
  world: "acme-billing",
  tasks: users.map((user) => taskSpecFor(user)),
  concurrency: 4,
  agent: async ({ toolkit, instruction }) => runCustomerAgent(toolkit.impls, instruction),
});

for (const result of report.results) console.log(result.task, result.reward, result.error ?? "");
```

<Info>
  `runSessions` hands its handler a bound toolkit. An agent that needs an HTTP
  base URL opens its own sessions with
  `openSession(world, { task, surfaces: ["api"] })` in a pool. See
  [Run many sessions at once](/worlds/parallel).
</Info>

## Stage 6: read the results per user

**Per run.** Each `grade` answers for one session, and each `SessionRunResult` carries `task`, `sessionId`, `reward`, `rewards` and an `error` when the run threw instead of grading. Keep the user id in the task's `metadata` to attribute the reward.

**Per user, later.** `gateway worlds task list --world acme-billing --metadata externalUserId=u_42` finds every stored task filed against one of your users.

**In the dashboard.** A world's own **Results**, **Runs** and **Evals** tabs hold what ran against it, **Overview** lists its live sessions with engine, state, task, API URL and age, and the Worlds section's **All results** tab is the matrix across every world in the project.

## Look a user up in a hashed world

A hashed field does not answer a plaintext filter. `--resolve` turns each value you pass into the form the store actually holds, per the world's `[redact]` policy, and runs the query on that.

```bash theme={null}
gateway worlds describe acme-crm
gateway worlds data query acme-crm accounts --where owner_email=jo@acme.test --resolve
gateway worlds data query acme-crm accounts --where plan=pro --limit 50
```

`worlds describe` stamps every field with its treatment — `[hash]`, `[preserve]`, `[round:n]`, `[drop]` — so you can tell before writing a query which fields will answer. A dropped field resolves to nothing and matches no row; the `resolved <field> -> …` lines on stderr say so.

Rows go to stdout as one compact JSON object per line, so the output pipes into `jq`; the `n of total` footer goes to stderr. The same query is `GET /api/public/worlds/{slug}/data?entity=&where=&resolve=true` and the `query_world_data` MCP tool.

<Info>
  A hash is salted with the connector's slug from `connector.toml`, not the
  world's platform slug, so the same address is a different digest in two
  worlds. A cross-world lookup recomputes both forms from the plaintext. See
  [Link worlds together](/worlds/links).
</Info>

## Before you hand it to customers

**Prove the seal.** Open a session on two tenants' worlds and read the state of each; neither should hold the other's records.

```bash theme={null}
gateway worlds session state "$ACME_SESSION"
gateway worlds session state "$GLOBEX_SESSION"
```

**Check what actually landed.** `gateway worlds data show acme-crm` prints the world's entity counts, and `--entity accounts` prints that entity's rows.

**Re-read one redacted field.** Query it with `--resolve` and without; the plaintext form should match only through resolution.

**Close what you open.** A loop that opens sessions and never closes them holds containers until the idle reaper finds them thirty minutes later, and a pinned session is never reaped at all.

## Where to go next

* [Simulations for your users](/worlds/for-your-users) for the provisioning API and what gets attributed to a key.
* [Put data in a world](/worlds/data) for the rows file, batches and million-row uploads.
* [Keep real data out](/worlds/redaction) for the policy, the salt and the four modes.
* [Evals in CI/CD](/use-cases/evals-in-ci) for the other recipe.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.