Skip to main content
Claude and OpenAI’s models ship their own computer-use tool: the model looks at a screenshot and answers with clicks, typing and key presses at pixel coordinates. native_computer serves that tool from a world session’s browser, so a model drives the world’s UI exactly as it would a desktop, and the grader scores what the clicks did. It is the pixel twin of the browser tools, which act by selector.

Open one

Open the session with the ui and browser surfaces, open the browser, then ask the session for a native computer for your provider.

Claude, from Python

execute takes a response’s content, runs every computer action in it in order and returns the tool_result blocks for the next user turn.

OpenAI, from Python

execute takes one computer_call, runs its batch of actions and returns the computer_call_output with the screenshot after them.
A computer_call with pending_safety_checks is not run: execute raises PendingSafetyChecks with the checks. Show them to a person, then call computer.execute(call, acknowledged=error.checks) to run it and acknowledge them.
An agent working across several worlds can hold one browser per world at once, in the same thread: open session.browser() on each world’s session and close each one when done.

From TypeScript, with the Vercel AI SDK

computer.aiSdk() is the options object for the AI SDK’s provider-defined tools: execute (and toModelOutput or needsApproval) backed by the session’s browser. Tool calls from one step run one at a time, in order.
The raw APIs work from TypeScript too: await claude.execute(response.content) and await gpt.execute(call, acknowledged?) return the same shapes as the Python examples. With OpenAI’s safety checks, aiSdk() sets needsApproval, so the AI SDK asks for approval before the call runs.

Any other agent framework

computer.execute_action(action) (TypeScript executeAction) runs one action in the provider’s own shape and returns what it produced (image bytes for a screenshot or zoom, text for the cursor position). For the OpenAI Agents SDK, AgentsComputer is a ready computer, and computer.tool() is its ComputerTool:
The Agents SDK calls its computer from an event loop and Python’s Playwright is synchronous, so AgentsComputer opens the browser on a worker thread of its own and runs every action there. computer.tool() always sets the tool’s safety-check gate: a call the model flagged runs only when on_safety_check (given {"id", "code", "message"}, sync or async) answers exactly True, and is refused otherwise or without one. A bare ComputerTool(computer=computer) would skip that gate, so use computer.tool().

What to know

  • Coordinates are the screenshot’s. The session browser is 1280x800, within both providers’ image limits, so nothing is scaled. Pass display=(w, h) (TypeScript { display: [w, h] }) to show the model a smaller screen: screenshots shrink and clicks are mapped back to the page. The display keeps the viewport’s shape (for 1280x800, 1024x640 works and 1024x768 is refused).
  • Malformed actions are refused. A missing coordinate or text, or an unknown action or button, is refused before it touches the page. With OpenAI the whole computer_call is checked before any of it runs; with Claude, each action answers for itself and the rest of the turn is not executed.
  • A turn stops at the first failure. With Claude, the failing action answers is_error with the error, and every later action in that turn answers Not executed: an earlier computer action in this turn failed.
  • Keys. Claude’s ctrl+a, Return, alt+Tab and OpenAI’s ["CTRL", "A"], ["ENTER"] both map to the browser’s key names.
  • On the record. Every action lands on the session’s browser timeline as computer_click, computer_type and so on. Text typed into a password field, one named like a token, key or PIN, a frame from another origin, or anything that is not a text field (a closed shadow root, the page body) is recorded as [redacted], and so is text shaped like an API key.
  • Traced with screenshots. When tracing is on, each action is a computer.<action> span whose output shows the screen after it. See Screenshots on spans.

Where to go next