> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfacearea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Native computer use (Anthropic, OpenAI)

> Hand a world session's browser to Claude's or OpenAI's own computer-use tool. The model sees screenshots and clicks at pixel coordinates; the SDK runs each action on the session's browser, returns the screen in the shape the provider expects, and traces every step with its screenshot.

Claude and OpenAI's models ship their own computer-use tool: the model looks at a screenshot and
answers with clicks, typing and key presses at pixel coordinates. `native_computer` serves that tool
from a world session's browser, so a model drives the world's UI exactly as it would a desktop, and
the grader scores what the clicks did.

It is the pixel twin of [the browser tools](/sdk-ts/browser#hand-the-browser-to-a-model), which act
by selector.

## Open one

Open the session with the `ui` and `browser` surfaces, open the browser, then ask the session for a
native computer for your provider.

| Provider | Kind | Tool sent to the model | Models |
| - | - | - | - |
| Anthropic | `"anthropic"` | `{"type": "computer_toolset_20260801"}` (no beta header) | `claude-sonnet-5-5`, `claude-opus-5-5` and the other Claude 5 models |
| Anthropic | `"anthropic-computer-20251124"` | `computer_20251124` (beta `computer-use-2025-11-24`) | earlier models; Claude 5.5 on Amazon Bedrock |
| OpenAI | `"openai"` | `{"type": "computer"}` (Responses API) | `gpt-5.5`, `gpt-5.6-sol` and the other models whose page lists computer use |
| OpenAI | `"openai-preview"` | `computer_use_preview` | `computer-use-preview` |

## Claude, from Python

`execute` takes a response's `content`, runs every computer action in it in order and returns the
`tool_result` blocks for the next user turn.

```python theme={null}
import anthropic
from gatewaysdk.world_sessions import open_session

client = anthropic.Anthropic()

with open_session("acme-billing", "refund-double-charge", surfaces=["ui", "browser"]) as session:
    session.ready()
    with session.browser() as browser:
        browser.goto()
        computer = session.native_computer(browser, "anthropic")
        messages = [{"role": "user", "content": session.instruction}]
        while True:
            response = client.messages.create(
                model="claude-sonnet-5-5", max_tokens=32000, tools=[computer.tool()], messages=messages
            )
            messages.append({"role": "assistant", "content": response.content})
            if response.stop_reason != "tool_use":
                break
            messages.append({"role": "user", "content": computer.execute(response.content)})
    print(session.grade())
```

## OpenAI, from Python

`execute` takes one `computer_call`, runs its batch of actions and returns the
`computer_call_output` with the screenshot after them.

```python theme={null}
from openai import OpenAI
from gatewaysdk.world_sessions import open_session

client = OpenAI()

with open_session("acme-billing", "refund-double-charge", surfaces=["ui", "browser"]) as session:
    session.ready()
    with session.browser() as browser:
        browser.goto()
        computer = session.native_computer(browser, "openai")
        response = client.responses.create(model="gpt-5.6-sol", tools=[computer.tool()], input=session.instruction)
        while calls := [item for item in response.output if item.type == "computer_call"]:
            outputs = [computer.execute(call) for call in calls]
            response = client.responses.create(
                model="gpt-5.6-sol", tools=[computer.tool()], previous_response_id=response.id, input=outputs
            )
    print(session.grade())
```

<Info>
  A `computer_call` with `pending_safety_checks` is not run: `execute` raises `PendingSafetyChecks`
  with the checks. Show them to a person, then call `computer.execute(call, acknowledged=error.checks)`
  to run it and acknowledge them.
</Info>

An agent working across several worlds can hold one browser per world at once, in the same thread:
open `session.browser()` on each world's session and close each one when done.

## From TypeScript, with the Vercel AI SDK

`computer.aiSdk()` is the options object for the AI SDK's provider-defined tools: `execute` (and
`toModelOutput` or `needsApproval`) backed by the session's browser. Tool calls from one step run one
at a time, in order.

```typescript theme={null}
import { anthropic } from "@ai-sdk/anthropic";
import { openai } from "@ai-sdk/openai";
import { generateText, stepCountIs } from "ai";
import { openSession } from "@withgateway/sdk/worlds";

const session = await openSession("acme-billing", { task: "refund-double-charge", surfaces: ["ui", "browser"] });
await session.ready();
const browser = await session.browser();
await browser.goto();

const claude = session.nativeComputer(browser, "anthropic");
await generateText({
  model: anthropic("claude-sonnet-5-5"),
  tools: { computer: anthropic.tools.computerToolset_20260801(claude.aiSdk()) },
  prompt: session.instruction ?? "",
  stopWhen: stepCountIs(50),
});

const gpt = session.nativeComputer(browser, "openai");
await generateText({
  model: openai.responses("gpt-5.6-sol"),
  tools: { computer: openai.tools.computer(gpt.aiSdk()) },
  prompt: session.instruction ?? "",
  stopWhen: stepCountIs(50),
});

console.log(await session.grade());
await browser.close();
await session.close();
```

The raw APIs work from TypeScript too: `await claude.execute(response.content)` and
`await gpt.execute(call, acknowledged?)` return the same shapes as the Python examples. With
OpenAI's safety checks, `aiSdk()` sets `needsApproval`, so the AI SDK asks for approval before the
call runs.

## Any other agent framework

`computer.execute_action(action)` (TypeScript `executeAction`) runs one action in the provider's own
shape and returns what it produced (`image` bytes for a screenshot or zoom, `text` for the cursor
position). For the OpenAI Agents SDK, `AgentsComputer` is a ready computer, and
`computer.tool()` is its `ComputerTool`:

```python theme={null}
from agents import Agent, Runner
from gatewaysdk.native_computer.agents_sdk import AgentsComputer   # pip install 'gatewaysdk[openai-agents]'

def ask_a_person(check: dict) -> bool:
    return input(f"{check.get('message')} Run it anyway? [y/N] ") == "y"

computer = AgentsComputer(lambda: session.browser(), on_safety_check=ask_a_person)
agent = Agent(name="clerk", model="gpt-5.6-sol", tools=[computer.tool()])
result = await Runner.run(agent, session.instruction)
computer.close()
```

The Agents SDK calls its computer from an event loop and Python's Playwright is synchronous, so
`AgentsComputer` opens the browser on a worker thread of its own and runs every action there.
`computer.tool()` always sets the tool's safety-check gate: a call the model flagged runs only when
`on_safety_check` (given `{"id", "code", "message"}`, sync or async) answers exactly `True`, and is
refused otherwise or without one.
A bare `ComputerTool(computer=computer)` would skip that gate, so use `computer.tool()`.

## What to know

* **Coordinates are the screenshot's.** The session browser is 1280x800, within both providers'
  image limits, so nothing is scaled. Pass `display=(w, h)` (TypeScript `{ display: [w, h] }`) to
  show the model a smaller screen: screenshots shrink and clicks are mapped back to the page. The
  display keeps the viewport's shape (for 1280x800, `1024x640` works and `1024x768` is refused).
* **Malformed actions are refused.** A missing coordinate or text, or an unknown action or
  button, is refused before it touches the page. With OpenAI the whole `computer_call` is checked before any
  of it runs; with Claude, each action answers for itself and the rest of the turn is not executed.
* **A turn stops at the first failure.** With Claude, the failing action answers `is_error` with the
  error, and every later action in that turn answers
  `Not executed: an earlier computer action in this turn failed.`
* **Keys.** Claude's `ctrl+a`, `Return`, `alt+Tab` and OpenAI's `["CTRL", "A"]`, `["ENTER"]` both
  map to the browser's key names.
* **On the record.** Every action lands on the session's browser timeline as `computer_click`,
  `computer_type` and so on. Text typed into a password field, one named like a token, key or PIN,
  a frame from another origin, or anything that is not a text field (a closed shadow root, the page
  body) is recorded as `[redacted]`, and so is text shaped like an API key.
* **Traced with screenshots.** When [tracing](/tracing/setup) is on, each action is a
  `computer.<action>` span whose output shows the screen after it. See
  [Screenshots on spans](/tracing/metadata#screenshots-on-spans).

## Where to go next

* [Give a world a UI](/worlds/ui) for the page the browser opens.
* [Drive a world UI](/sdk-ts/browser) for the browser itself and selector-based tools.
* [Spin worlds up and down](/worlds/sessions) for the session around it.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.