
Why a world, and not a test suite
A test suite checks the responses your agent produced. A world checks the decisions it makes when the system underneath it behaves the way the real one does. Some behaviour cannot be written down as an input and an expected output:- A refund is denied because the ledger says the charge already settled.
- A sanctions screen half-matches a name at 0.91.
- An upstream service starts timing out on the third call.
If you can write the expected answer down in advance, you want an eval. If the
answer depends on what the systems around the agent do, you want a world.
One world per system, not one per test
A world models a system your agent has to deal with, not a single case. Build one world for checkout, one for support triage, one for refunds, then ask each of them many questions. One world holds many scenarios, many rollouts, and many runs over time. Make a new world only when the agent starts working against a genuinely different system. If two situations call the same services and read the same data, they belong in the same world.Sealed, and pinned to a moment
A world starts from captured sessions, selected connector data, or an explicitly labeled synthetic workspace. Two properties keep the starting state useful across repeated runs. Sealed means world tools change only the world’s stored state. Captured or synthetic records belong to the replica, and tool writes do not change the connected customer system. Pinned means every world has a version, and every run records the version it ran against. Without the version a pass rate is uninterpretable: you would not know whether the number moved because the agent changed or because the world did.Pinning is what makes two runs comparable. When a run is recorded as
0.8.0 · 939cd2df, the second half is the exact world state it faced. Two runs against
the same pinned state, with the model held fixed, differ by exactly the thing
you changed.One session becomes many questions
For a populated collaboration example, build a Slack world from Connectors. Choose a deterministic synthetic workspace or capture selected public/private conversations, then run an agent against the published version and grade its replies and reactions. Captured sessions are the seed. Replay only tests the path that already happened. A world also generates the situations that did not:Failures, in both directions
Reactively. Promote a production failure into a scenario. Every candidate from then on has to get it right before it can ship. Proactively. Apply the shapes in the table above to a workflow that has never failed, and find where it breaks before a customer does.Comparing models against your own systems
Because the world is pinned, everything except the model can be held fixed. Hold the agent, tools, policy, and cases constant, change only the model, and read the runs against each other. The answer is specific to your systems.Where to go next
Getting started builds one world end to end: contract, data, version, session, graded run. Run a world. Driving a world that already exists.
Author a world. Building one and keeping it honest.
Elsewhere on the site.
- Glossary for the six words the whole product is built from.
- Tracing to capture the sessions a world is seeded with.
- Evaluation & Replay for how a world’s runs are scored.
- Worlds client to inspect worlds, scenarios, and rollouts from Python.