> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Tool calls into rows

> Record an agent's real tool calls with captureToolCalls, then put them into a world with ingest - through your own transform, a declared ingest.toml mapping, or the operation's projection - and the world's contract. Redaction optional, off by default.

A tool's response is not the entity's shape, and nothing is inferred. `@withgateway/sdk/worlds`
does two things: records the calls while the agent runs against the real vendor, and puts the
records into a world through the route you declare — your own `transform`, an `ingest.toml`
mapping, or, for a tool whose result is the vendor's response body, the operation's projection.

```ts theme={null}
import { captureToolCalls, ingest, readToolCallRecords } from "@withgateway/sdk/worlds";

// 1. The agent's tools, exactly as before, every call recorded to a file.
const tools = captureToolCalls(
  {
    list_users: async (args) => vendor.listUsers(args),
    get_account: async (args) => vendor.getAccount(args),
  },
  "calls.jsonl",
);
await runMyAgent(tools);

// 2. The records into the world, shaped by your code. Nothing is redacted unless you ask.
const report = await ingest("acme-crm", readToolCallRecords(["calls.jsonl"]), {
  transform: (records) => ({
    users: records.filter((r) => r.tool === "list_users").flatMap((r) => r.result.users),
  }),
});
console.log(report.rows, report.refusals);
```

## The record

One JSON object per call. `tool`, `args` and `result` are required; the rest is optional and
travels with the record.

```json theme={null}
{
  "tool": "list_users",
  "args": { "team": "t1" },
  "result": { "users": [{ "id": "U1", "email": "a@x.test" }] },
  "at": "2026-09-15T10:00:00Z",
  "session": "optional",
  "source": "optional free text"
}
```

A line of `gateway worlds session export` (`{seq, kind, tool, args, result, error, completedAt}`)
is accepted as it is. `parseToolCallRecord(raw)` tells you what the runtime will do with a
record before you send it: `{ record, skip, refusal }` — a `seed` line or a failed call is a
`skip` with its reason, a record with no tool name or no result is a `refusal`.

## `captureToolCalls(tools, sink, options?)`

Wraps a `name → async function` map — the shape the Vercel AI SDK's tools and a
`WorldToolkit`'s `impls` share — and returns the same map. Each wrapped function calls the
original, returns exactly what it returned, and writes one record to the sink. An error is
thrown through untouched and recorded with `error` instead of `result`, which ingest later
skips and counts. The agent changes nothing.

| Sink | What happens |
| - | - |
| an array | the record is pushed |
| a file path | appended as one JSON Lines record |
| a function | called with the record and awaited |

`options.session` and `options.source` are stamped on every record.

## `ingest(world, records, options?)`

`world` is a local directory or a platform slug (`slug` or `slug@ref`). `records` are tool-call
records, or any JSON objects when a mapping reads them as `rows`.

| Option | Default | Meaning |
| - | - | - |
| `transform` | — | Your `transform(records) => {entity: [rows]}`: a function, run here; or a script path (`.py`, `.ts`, `.mjs`), run in a subprocess with a time limit (`timeoutMs`, 120 000), sockets disabled and no `GATEWAY_*` variables. Its stderr is in the report. |
| `map` | — | An `ingest.toml` mapping: its text, or the parsed table (`IngestMapping`). Exclusive with `transform`. The grammar is on [Put data in a world](/worlds/data#ingest-from-any-source-a-declared-mapping). |
| `redact` | `"off"` | Optional. `"apply"` runs the world's `[redact]` policy over the rows first; `"refuse"` rejects plaintext in a redacted field. |
| `dryRun` | `false` | Produce and prove the rows; write and upload nothing. |
| `mode` | `"append"` | `"append"` upserts by primary key; `"replace"` replaces the named entities. |
| `message` | — | Change reason for the version a slug cuts. |
| `host`, `publicKey`, `secretKey` | the environment | As on every worlds helper. |

With neither `transform` nor `map`, the world's own `ingest.toml` applies when it has one;
else each record's `tool` resolves to an `[[operations]]` entry of `connector.toml` and its
`result` is read at the operation's `results` path and projected by its `[ingest]` table
(`drop`, `explode`, `carry_args` from the call's `args`) as a [captured response](/worlds/custom-connections)
is. The report says `via: "operations"`, and its `note` says that this is the operation's
projection, not a mapping. `ingestToolCalls(world, records, options)` is that route by name.

* A **directory** is written in place (`data/initial.json`, the snapshot rebuilt) through the
  bundled schema runtime, which needs a `python3` ≥ 3.12 on the machine. Mappings and
  projections run there; a transform runs where you are.
* A **slug** is pulled to a temporary directory, the rows produced and proven there, and
  pushed through the chunked data batch lane as a new version.

The report:

```ts theme={null}
type IngestReport = {
  world: string;
  via?: "transform" | "map" | "operations" | "rows";
  records: number;                   // records given
  ingested: number;                  // records that yielded rows
  skipped: Record<string, number>;   // by reason: "kind seed", "call failed", "operation ping lands no entity"
  refusals: { index: number | null; tool: string | null; reason: string }[];
  rows: Record<string, number>;      // rows produced, per entity
  sources?: {                        // via map: one per [[sources]]
    name: string; kind: string; select: string; entity: string;
    records: number; items: number; rows: Record<string, number>;
    unmappedFields: Record<string, string>;   // source fields nothing reads, with a sample
    unsetFields: Record<string, string[]>;    // entity fields never set
  }[];
  unmatched?: Record<string, number>; // via map with unmatched = "skip": records no source claimed
  note?: string;                     // via operations
  stderr?: string;                   // via transform: what the script wrote
  dryRun: boolean;
  redact: "off" | "apply" | "refuse";
  changed: boolean;
  entityCounts?: Record<string, number>;
  batch?: Record<string, unknown>;   // a pushed slug: the batch's final status, with versionId
};
```

A record no mapping source claims, or a tool that is not an operation, is a refusal naming
it, and with any refusal nothing is written. A row the world's contract refuses throws, as
`data import` would, naming the entity, index and field. A transform that fails, times out or
returns something other than `{entity: [rows]}` throws with its stderr.

The mapping as an object:

```ts theme={null}
import type { IngestMapping } from "@withgateway/sdk/worlds";

const map: IngestMapping = {
  ingest: { on_conflict: "latest", unmatched: "skip" },
  sources: [
    {
      name: "crm_read",
      select: "result",
      when: { tier: { present: true } },
      entity: "accounts",
      fields: { id: "id", name: { from: "name", transform: ["trim"] }, tier: { from: "tier", default: "smb" } },
      explode: [{ path: "contacts", entity: "contacts", carry: { account_id: "id" }, fields: { id: "id", email: "email" } }],
    },
    { name: "crm_log_activity", select: "args", entity: "activities", fields: { id: "$source.id", note: "note", at: "$at" } },
  ],
};
await ingest("acme-crm", records, { map, dryRun: true });
```

<Info>
  Redaction is optional and off by default: rows land as they were produced. Use
  `redact: "apply"` for a world built from real users. See
  [Keep real data out](/worlds/redaction).
</Info>

## From traces

Instrumented agents already record every tool call as a TOOL observation. Convert those
instead of intercepting:

```ts theme={null}
import { sessionToolObservations, toolCallsFromObservations, ingestToolCalls } from "@withgateway/sdk/worlds";

const observations = await sessionToolObservations("<sessionId>");
const { records } = toolCallsFromObservations(observations, { session: "<sessionId>" });
await ingest("acme-crm", records, { transform: "./shape.ts" });
```

`toolCallsFromObservations` takes `name` as the tool, `input` as the args, `output` as the
result and `startTime` as the time, parses serialized JSON back, and sets `source` to
`trace:<traceId>/<observationId>`. Only `TOOL` observations convert; an `ERROR`-level one
becomes a failed call. With `tools` (the world's operations) every other tool's observation
lands in `unknownTools` instead — the filter the operations route uses. The CLI form is
`gateway worlds data calls pull <out> --from traces`, then `gateway worlds data ingest`; see
[Put data in a world](/worlds/data#pull-transform-ingest).

## Python

The same two halves ship in `gatewaysdk`:

```python theme={null}
from gatewaysdk.world_tools import capture_tool_calls
from gatewaysdk.world_data import ingest_records

tools = capture_tool_calls({"list_users": vendor.list_users}, "calls.jsonl")
run_my_agent(tools)

def transform(records):
    return {"users": [u for r in records if r["tool"] == "list_users" for u in r["result"]["users"]]}

report = ingest_records("acme-crm", records, transform=transform)            # a callable, or a script path
report = ingest_records("acme-crm", records, mapping="ingest.toml", dry_run=True)   # a mapping: path, text or table
```

`ingest_tool_calls(world, records)` is the operations route by name. Both take `redact`
(`"off"` by default), `dry_run`, `mode`, `message` and the keys.

## Where to go next

* [Put data in a world](/worlds/data) for `calls pull`, the transform, the mapping grammar and the batch lane.
* [Simulations from real data](/use-cases/sims-from-real-data) for the whole funnel.
* [Keep real data out](/worlds/redaction) for the `[redact]` policy and the modes.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.