Skip to main content
A session is one live container of a world, held warm while your agent works against it. The agent stays in your process. Every call() executes inside the platform’s container and is recorded in the session’s call log. Sessions show up under Sessions on the platform without your agent instrumenting anything.

Open a session

openSession(slugOrRef, { task, versionId?, reuseWarm?, surfaces?, browserTier?, options? }) returns as soon as the platform accepts the session. Call ready() to wait for the container.
world.open(task, opts?) on a WorldSet does the same thing and pins the session to the version the world already resolved. Asking for a surface the world does not declare is refused with HTTP 400, rather than served with fewer surfaces.

Attach to a session someone else opened

attachSession(sessionId, options?) returns a WorldSession for an existing session. Use it to resume after a crash, or to pick up a session that CI opened and handed to a developer.

Read what the session is

These are properties, not requests. They reflect the last state the session read.

Wait for the container

ready(opts?) polls until the session answers calls: the World Host holds it, or the process inside a container has asked the platform for work. It waits ten minutes by default and throws WorldSessionTimeout past the deadline, or WorldSessionOpenFailed (a WorldSessionCallError carrying the platform’s reason and code) if the session fails or stops first.
status() reads the session’s latest state in one request and makes no attempt to wait.

Call the world’s tools

call(tool, args?, opts?) runs one of the world’s tools inside the container. The arguments and the result are the world’s own contract. Nothing in between validates or reshapes them, so a tool error reads exactly as it would in a hosted run.
A call fails with WorldSessionCallError when the tool errors, and WorldSessionTimeout when the container does not answer inside three minutes.

Hand the world’s tools to a model

toolkit() returns the world’s tools in the shape a model expects. Schemas come from the world’s own signatures, and each implementation calls into this session.
A tool added to the world appears here on the next run. A tool renamed in the world renames here.

Point an existing client at the world

api() returns the base URL of the world’s own HTTP API for this session. apiHeaders() returns the headers every request to it needs.
A World Host session gates its API surface on a per-session token, and apiHeaders() returns it as { Authorization: "Bearer <token>" }. A pod session or a customer-hosted twin answers with the world’s own declared auth, and apiHeaders() returns {}.

Send the token the way your client already sends credentials

Worlds whose vendor uses a header or basic scheme accept the same token the vendor’s own way. Read it from the surfaces block and build whichever header your client already sends.
Pointing a client at a world is a base URL and a key. No code in your client has to know a world exists.
Send the token as a header, never in a URL. A query token outlives the request in access logs, and the platform refuses it.
surfaces() returns the whole block, and ui() returns the URL of the world’s UI entry page with the session token already attached. Both throw WorldSurfaceUnavailable when the session was not opened asking for that surface.

Put rows in, read rows back, start over

seed(rows, opts?) writes rows through the world’s contract in one batch. append is the default and upserts by primary key. replace makes the named entities hold exactly these rows. The batch is atomic, so a refused row leaves the session unchanged and the error names the entity, row and field.
Seeded rows are stored as given. Pass { redact: "apply" } to run the world’s [redact] policy over them first, or { redact: "refuse" } to reject a row carrying plaintext in a redacted field; both need the world to have a connector.toml. See Keep real data out. state(opts?) returns every entity’s rows as the world holds them now, which is the same dump a grader reads.
reset(opts?) returns the world to its opening state: the version’s data plus the task’s seed. Rows written since are gone and the call log stays.

Grade the end state

grade(opts?) runs the task’s grader against the container’s end state and returns { reward, rewards, raw }. reward is the scalar, rewards holds every named component, and raw is the grader’s untouched answer.
A rubric grade is finalized by a model on the platform rather than by the container. The call then waits on the judge’s budget of fifteen minutes instead of a call’s three. Pass timeoutMs to override either.

Trace your agent inside the session

withTraceContext(fn) runs fn as work inside this session. Every span it records takes the session’s id as its session id and carries world_session_id in its trace metadata, so your agent’s traces sit beside the platform’s trace of the session’s tool calls. runSessions wraps every agent call in it.
A tracing.withSession(...) block inside still sets the session id. Nothing is recorded until tracing.init() has run.

Read the episode back

export() returns the session’s call log oldest first, paging through the whole thing. Every tool call, seed, reset and grade is in it, with its arguments, result and error.
export() holds what the platform brokered. The world’s own request log is the calls a grader reads: every call the world answered, its route requests (including the ones your agent made through the browser or straight to api.url) and its tools called by name, each with its via and with credential fields read as [redacted]. WorldRequestLog.read(session) reads it from the live session, so read it before you close.

Keep an inline task

saveTask(opts?) promotes this session’s inline task into a stored task of the project, at version 1, pinned to the session’s world.
A session on a bundled task has nothing to save. A session on a stored task is already saved.

Close it

close(opts?) ends the session. keepWarm: true hands the container back to the pool so the next session on the same task starts warm.
Always close in a finally block. A leaked session holds its container until it closes for being idle.

Write the task yourself

Instead of a bundled scenario name, pass an inline WorldTaskSpec. It is one end user’s task: what the agent is told, the rows the session starts from, and how it is graded.
Without a grader, grade() answers reward: null.

The three grader kinds

A check’s assert is exists by default, and may be absent, count, all or any. Each check’s id names its entry in rewards, and the task’s reward is the weighted share of checks that passed. A python grader reads the end state as JSON on stdin: each entity as { "<primary key>": row }, the call log as calls, and the agent’s report as report (a cross-world verifier sees <alias>.<entity>). It prints its grade as the last line of stdout, one flat JSON object: reward plus one key per check, each a number from 0 to 1. true and false read as 1 and 0, and null or a string marks that check ungraded. Without reward the task is ungraded. The script must exit 0 within timeoutSeconds (default and maximum 30). A nested object or array, a number outside 0 to 1, a non-zero exit, a last line that is not JSON, or a timeout fails the grade with the reason.
Inline and stored tasks run on a host engine world.

Store a task so many sessions can open it

An inline task belongs to one session. A stored task is the same spec as a named, versioned project resource. Updating the spec appends a version rather than rewriting history, and a session opened on { id, version: 2 } keeps running version 2 after version 3 exists.
world pins a task to one world’s slug, and the spec is validated against that world’s contract. Passing null leaves the task usable with any world.
validateTask reports checked: "entities" when the world’s contract was read and every seed and check entity was matched against it. It reports checked: "none" when the version has no contract to check, in which case ok says only that the version exists and is ready. Tasks are project-scoped. Another project’s task id comes back as not found, never as forbidden.

Tasks that need several worlds

A task whose spec has a worlds table is one instruction over several hosted worlds. openTask brings every one of them up in one platform call and hands back the sessions keyed by the aliases the spec chose.
Each entry of worlds is { slug, ref?, seed?, grader?, tools? }: the world, one of its refs (default main), and that world’s share of the task. With worlds present the top-level seed, grader and tools are refused and the task is not pinned to a world. The instruction and metadata are shared. TaskSessions carries taskId, taskVersion, name, instruction, sessions (alias to WorldSession) and aliases, plus the task-wide verbs, each fanned out over every world: Open a multi-world task with openTask, not openSession. Under the hood openTask is POST /api/public/world-task-sessions and listTaskSessions is GET /api/public/world-task-sessions?taskId=.

Run a task end to end

runTask is gateway worlds task run as a function: open every world of a task, run your agent once against all of them, grade with its own report, file one rollout and close — the TypeScript twin of the Python SDK’s gatewaysdk.run_task.
The agent is called as agent(run) with a TaskRun: The agent’s return value is the report every world’s grader sees: a string, or an object with a report key (any other keys are kept in run.json as-is); a Promise is awaited. runTask’s options: Every step happens even when the agent throws: it still grades (with report: null), still files (the evaluation’s status FAILED, run.json.agentError set to the error’s message), and still closes every session unless keepUp: true. A world whose tools the platform could not read (a tool-contract-unread notice) throws WorldToolsUnread (its code is tool-contract-unread) before the agent runs: every session closes, even with keepUp: true, and nothing is filed. runSessions records such a task as failed without calling its agent. The run writes <outDir>/<runId>/{manifest,grades,run}.json (and report.md when the agent gave a report) and returns the same object as run.json: { runId, taskId, taskVersion, name, instruction, model, aliases, reward, rewards, ungraded, agentError, report, traced, runDir, durationSeconds, notes, filed, evaluationId? }.

Group a task run’s own spans into one session

tracing.withSession(sessionId, fn) groups every trace recorded inside fn under one session id, the TypeScript twin of the Python SDK’s tracing.session(session_id). runTask uses it to put the agent’s own LLM calls and tool spans under a session named by the run id:
withSession(sessionId, fn) sets the session id in an AsyncLocalStorage for the duration of fn (so it survives awaits and propagates to auto-instrumented spans), and the SessionSpanProcessor — registered automatically by tracing.init()/initIsolated() — stamps gateway.session.id on every span started while it is set. Nesting is fine: the innermost withSession() wins for the code that runs inside it. A gateway.session.id the span is started with, or set on it afterwards with span.setAttribute, overrides this ambient default for that span.

Where to go next