Skip to main content
Hand each of your customers a sealed copy of their own systems to run agents against. The data comes from traffic you already have, redaction on the way in keeps personal fields out, and every run is graded against the state the agent left behind. Six stages, each one a command.
Decide the redaction policy before the first import, not after. Rows already in a world were stored as given unless you asked otherwise, and re-hashing a digest that is already hashed breaks every join that depended on it.

Stage 1: get the real rows out

Three sources.

Your agent’s own runs, over OTLP

Instrument the agent with the gatewaysdk Python package and every LLM call, tool call and nested step lands in your project as a trace, grouped into sessions. Traces show which situations to build: the requests real users made, and what the vendor answered.
Every tool call in those traces is a TOOL observation — the tool’s name, its input and its output. A tool’s output is not the entity’s shape, so the route is pull, transform, ingest:
calls pull takes --trace (repeatable), --session or --since, and writes {tool, args, result, at, source} records. The transform is your transform(records) -> {entity: [rows]} in Python or TypeScript; it runs in a subprocess, and the rows then go through exactly what data import does. --map ingest.toml is the no-code alternative, a declared mapping in which nothing is inferred. See Pull, transform, ingest. To intercept the calls at the source instead, captureToolCalls from the TypeScript SDK (capture_tool_calls in Python) wraps the agent’s tool functions so every real call against the vendor is written as the same record, with no change to the agent; worlds data ingest then reads the file. Traces remain the place to read which situations to build — the dashboard and the REST API hold them, and Export to Surface Area covers the endpoint, batching and flushing.

A session’s own call log

A session that already ran records every call it served. worlds session export writes the log as JSON Lines, oldest first.
Each line carries {args, completedAt, error, kind, result, seq, tool}, where kind is call, grade, seed, reset or close. Reads carry rows in their result; writes carry the row the world should hold. worlds data ingest reads the file as it is, through the same transform or mapping.

A vendor export, shaped by data extract

worlds data extract reads a file and writes rows already shaped for data import, against the world’s own contract.
Extract drops null and empty fields, strips fields the entity does not declare and counts them, dedupes by primary key (--merge fill keeps the first row, --merge upsert the last), and runs the world’s [redact] policy over every row unless you pass --no-redact. Its stderr summary is {"read": N, "rows": M, "unknownFields": {…}, "redacted": {"dropped": n, "hashed": n}}.

Stage 2: redact on the way in (optional)

Redaction is optional. Every write path stores rows exactly as given unless you ask otherwise, so a world seeded from already-sanitized data, or from data that carries nothing personal, skips this stage. When the rows do carry personal fields, this stage hashes, drops or shapes them. The world’s connector.toml declares one [redact] block, and every write path reads the same one.
hash stores a salted sha256:<64 hex> digest, drop removes the key entirely, preserve stores a format-preserving digest so an account number still looks like one, and round keeps a number and drops only its precision. Keep real data out explains what each mode costs. Every write path takes --redact, and every one of them defaults to off.
apply runs the policy before the contract sees the row. refuse rejects a row that still carries plaintext in a redacted field, naming the entity, index and field. Both need the world to have a connector.toml; a world without one is refused up front rather than quietly storing plaintext.
Use apply for a world with preserve fields: a shaped digest looks like a real identifier by design.

Stage 3: pick the shape that matches what differs

Start with one world and per-user tasks. Move to a world per tenant only when customers need different data or a different contract: every world is a version history you then have to keep.

One world per tenant

--from takes a shipped template slug, workspace:<slug> for a connector your organization published, or a world directory. --overlay ./tenants/acme lays a directory of files over the source before the tree is compiled, which is how one base contract becomes each tenant’s variant. Each tenant’s world is stored as its schema tree, so one session serves both the world’s tools and — when the connector declares routes — the vendor’s HTTP paths. Slugs are unique within a project, so prefix them with the tenant and settle the naming scheme before the first customer: a collision answers 409 rather than overwriting anything.

Many tenants in one call

--batch takes a JSON Lines file (or - for standard input) of one world per line and prints one result line for each.
A line may also carry data for the seed rows, as {"rows": {"<entity>": [...]}}. The command exits 0 when every world was created and 2 when any line failed, so a partial batch is visible in the exit code rather than only in the output. Provisioning from your own backend instead is POST /api/public/worlds, the same request — see Simulations for your users.

One world, a task per end user

A task is an instruction, an optional seed and a grader. Supplying one at session open lets a single published version serve any number of end users, each with their own goal and starting rows.
The seed block carries its own redact mode, so every session the task opens seeds the same way. metadata is echoed back on the session; put your own user id there. Keep a task your users return to as a project resource rather than a request body:
task validate checks the seed and grader entities against the world’s contract without opening a session. It catches a grader asserting on an entity the world does not define.

Stage 4 and 5: seed, run, grade

1

Open a session for the run

The descriptor carries sessionId, ready, the surfaces it opened, and — with --surface api — an api.url and api.token.
2

Add this user's rows, if the task did not

Seeding changes that one live copy and no version. --mode replace swaps the entities the file names; the default --mode append upserts by primary key.
3

Hand the user's agent the URL and the token

Point it at api.url with Authorization: Bearer <api.token>, as it would be pointed at the real vendor. Nothing else in its code changes.
4

Grade the end state

The answer is {reward, rewards, raw} — the headline number, the per-check breakdown, and whatever the grader returned.
5

Keep the ledger, then close

A session holds a running container and closes itself after thirty idle minutes. Closing it yourself returns the container immediately.
To run many users at once, runSessions from the TypeScript SDK opens concurrency sessions at a time and mixes bundled names, stored task refs and inline specs freely, so one call can run a task per user:
runSessions hands its handler a bound toolkit. An agent that needs an HTTP base URL opens its own sessions with openSession(world, { task, surfaces: ["api"] }) in a pool. See Run many sessions at once.

Stage 6: read the results per user

Per run. Each grade answers for one session, and each SessionRunResult carries task, sessionId, reward, rewards and an error when the run threw instead of grading. Keep the user id in the task’s metadata to attribute the reward. Per user, later. gateway worlds task list --world acme-billing --metadata externalUserId=u_42 finds every stored task filed against one of your users. In the dashboard. A world’s own Results, Runs and Evals tabs hold what ran against it, Overview lists its live sessions with engine, state, task, API URL and age, and the Worlds section’s All results tab is the matrix across every world in the project.

Look a user up in a hashed world

A hashed field does not answer a plaintext filter. --resolve turns each value you pass into the form the store actually holds, per the world’s [redact] policy, and runs the query on that.
worlds describe stamps every field with its treatment — [hash], [preserve], [round:n], [drop] — so you can tell before writing a query which fields will answer. A dropped field resolves to nothing and matches no row; the resolved <field> -> … lines on stderr say so. Rows go to stdout as one compact JSON object per line, so the output pipes into jq; the n of total footer goes to stderr. The same query is GET /api/public/worlds/{slug}/data?entity=&where=&resolve=true and the query_world_data MCP tool.
A hash is salted with the connector’s slug from connector.toml, not the world’s platform slug, so the same address is a different digest in two worlds. A cross-world lookup recomputes both forms from the plaintext. See Link worlds together.

Before you hand it to customers

Prove the seal. Open a session on two tenants’ worlds and read the state of each; neither should hold the other’s records.
Check what actually landed. gateway worlds data show acme-crm prints the world’s entity counts, and --entity accounts prints that entity’s rows. Re-read one redacted field. Query it with --resolve and without; the plaintext form should match only through resolution. Close what you open. A loop that opens sessions and never closes them holds containers until the idle reaper finds them thirty minutes later, and a pinned session is never reaped at all.

Where to go next