Record a named session
Give a session a stable id at record time. That id is what you file into a task set and replay against later:One trace per turn
A session should read turn-by-turn: one trace per user message. With auto-instrumented frameworks each top-level agent invocation becomes one trace, so run one invocation per turn and carry the history yourself. Don’t push a whole conversation through a single call, or the session collapses into one giant trace:session.turns() maps 1:1 to the conversation, which
turn-level selectors and full-conversation replay depend on.
Load a session
load_session calls GET /api/public/sessions/{id}?includeIO=true: one
request returns the complete session. Credentials resolve from the same
environment variables tracing uses, or pass host=, public_key=,
secret_key= explicitly.
Work with the recording
input, output, model, level,
status_message, and timestamps, untruncated.
Dump a fixture
export_session tool, so human-made and agent-made
fixtures are interchangeable.
Replay suites from task sets
A task set on Surface Area is a regression corpus: each task stores a prompt, the expected behavior, and provenance to the recorded session that motivated it (metadata.sourceSessionId). Sessions get filed into task sets
by hand or from a failing rollout (promote_failure_to_scenario): no repo change needed to add a case.
Simulate the recorded world
The source session recorded every tool output, so replays are deterministic: your agent runs against the catalog/database/API responses as they were:{"args_json": "..."} wrappers
and agent-as-tool {"input": "..."} envelopes are unwrapped before
comparison, so recorded inputs match the args your live agent passes):
- local: tools you declared live (see below); never drift
- exact: normalized-JSON input equality
- exact-repeat: identical re-reads of a
(name, input)pair return the same recording again instead of exhausting it - ordinal: nth un-consumed recording for that name (robust to arg phrasing drift across models/prompts)
miss_policy:
"error"(default): raiseStubMissError, strict replay"reuse": recordings for that tool exhausted? Replay the last one again so research-heavy agents (more searches than the recording holds) finish the run. Never-recorded names still raise."passthrough": call yourfallback=lambda name, args: ...live
local=, visible to the contract check but never
counted as drift:
"error" policy does not hard-stop
those runs: a dead recording usually surfaces as a max-turns exception
instead. Catch both, then triage with case.check.
Grade with three-way triage
Contract invariants come from task metadata:
expectTools (must all be
called), mustNotCall (none may be called), and optional
outputExact: true for literal output matching. Semantic grading stays with
you: assert on case.expected, or score the replay trace with a platform
evaluator.
Grade against the ENTIRE session: 1:1 trace replay
Every case retains its full recorded session (case.recording): all
turns, all observations, untruncated, not just the prompt and expected
output. Two things build on that:
Re-drive the whole conversation. case.conversation() returns the
recorded user message of each turn (each turn was one trace); feed them to
your agent in order to replay the full session, not just turn one.
Exact mode. For a true 1:1 comparison, grade with mode="exact" (or set
metadata.replayMode: "exact" on the task):
result.trace_diff. Use it for deterministic/pinned agents (temperature 0,
workflow-style agents); for stochastic agents prefer contract mode, and run
exact mode to see how behavior shifted. Agents making parallel tool calls can
interleave order legitimately: the recorded order is the observation order of
the original trace.
You can also assert on anything in the recording directly:
Run the whole suite
record=False):
- each case’s replay becomes its own session named
<run>-<case>, traced under the internalgateway-replayenvironment and taggedreplay,taskSet:<name>,case:<name>,run:<runName>, hidden from the sessions page by default; filter on thetaskSet:<name>tag to find them - the run is filed on the task set as a dataset run (one item per case, with the PASS/FAIL/DRIFT status and the source session id in item metadata), listed on the Runs tab behind Edit simulations, so you can compare runs across agent versions and score them with evaluators
taskSet:<name> yourself is found by the same
tag filter: task sets work as test suites for non-environment rollouts too.
Your agent can declare any of these signatures; the runner adapts:
Pytest integration
Pin the corpus into your repo
Task set API
Manage suites programmatically (Basic auth with project API keys):
Task metadata drives grading:
expectTools, mustNotCall, outputExact,
and replayMode: "exact" for 1:1 grading, all optional.
Where to go next
- Worlds client — run the agent against a live world instead of a recording.
- Evaluation & Replay — the platform side of replay and scoring.
- Tracing sessions — how a session gets recorded in the first place.