> ## Documentation Index
> Fetch the complete documentation index at: https://chatcli.edilsonfreitas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Browser Automation (@browser)

> Drive a real local Chrome/Chromium over the DevTools Protocol — open a page, read it as text with numbered interactive elements, click, type, run JS, screenshot, and inspect console + network. The verification loop for web work. Keyless, zero new dependency.

The **`@browser`** tool drives a **real local Chrome/Chromium** browser so the agent can **see and interact with web pages** the way a person would: open a URL, read the rendered page as text, click and type into elements, run JavaScript, capture screenshots, and inspect the page's **console** and **network** activity.

It is the **verification loop** for web work — the agent builds or changes a frontend, then **actually looks at it** and debugs it from what the page logged and requested, instead of guessing from source.

<Info>
  `@browser` speaks the **Chrome DevTools Protocol (CDP)** directly over the websocket client ChatCLI already ships — **zero new dependency**, no driver, no downloaded runtime, no API key. It only needs a **locally installed Chromium-family browser** (Google Chrome, Chromium, Brave or Edge).
</Info>

***

## How it works

```text theme={"system"}
@browser open http://localhost:3000
        │
        ▼
launch a headless Chrome (first use) — reused for the whole session
        │
        ▼
attach over CDP (websocket) → navigate → wait for load
        │
        ▼
snapshot: title, URL, NUMBERED interactive elements, visible text
        │
        ▼
click [n] / type [n] "text" → snapshot again → console / network to debug
        │
        ▼
close (or it dies automatically when ChatCLI exits)
```

1. **Lazy launch** — the first `open` starts one browser; every later command reuses that **single stateful session**.
2. **CDP over websocket** — navigation, DOM reads, screenshots and event capture all flow over the DevTools websocket. No Selenium/Playwright, no external binary.
3. **Snapshot as text** — the page is rendered to model-friendly text with every interactive element **stamped and numbered** so the agent can address them precisely.
4. **Continuous capture** — console messages (errors included) and network responses are captured into bounded rings, so you can ask "what did the page log?" **after the fact**.
5. **Visible on demand** — the browser is headless by default, but the agent can put the window **on your screen** with `show` (or `open --visible`) whenever *you* need to act — log in, pick an account, solve a captcha — then take it back with `hide`. The switch relaunches the browser **on the same profile**, so cookies and logins survive it.
6. **Dies with the session** — a browser ChatCLI launched never outlives it; it is closed on CLI shutdown (or explicitly with `close`).

***

## Subcommands

Invoke with a JSON envelope `{cmd, args}` or the flat argv form. The model calls `@browser` automatically in `/agent` and `/coder` when it needs to look at a page.

| Subcommand | What it does | Key args |
| :- | :- | :- |
| `open` | Navigate to a URL (launches the browser on first use) and return a **page snapshot** | `url` (req.), `--visible` |
| `show` | Make the window **visible on your screen** (relaunches a headless session on the same profile — logins kept); optionally navigates | `url` (optional) |
| `hide` | Back to headless; the window disappears, the session and its cookies stay | — |
| `wait` | Block until the page reaches a state: URL contains, text contains, an element exists and/or the URL **changes** — the hand-off after `show` | `--url`, `--text`, `--selector`, `--changed`, `--timeout` (120s, max 600s) |
| `snapshot` | The current page as text: title, URL, numbered interactive elements, visible text | `--max N` (cap text length) |
| `click` | Click an element by its **`[n]` ref** from the last snapshot **or** a CSS selector | `target` (req.) |
| `type` | Type into an input; `--submit` presses Enter / submits the form | `target` (req.), `text`, `--submit` |
| `press` | Send a key to the focused element: `Enter`, `Tab`, `Escape`, `ArrowDown`, a character, or a chord like `Control+a` | `key` (req.) |
| `hover` | Move the mouse over an element (opens hover menus and tooltips) | `target` (req.) |
| `select` | Choose a `<select>` option by value or visible text | `target` (req.), `value` |
| `upload` | Attach local file(s) to an `<input type=file>` — the OS picker never opens under automation | `target` (req.), `file` / `files` |
| `scroll` | Move the viewport | `down` · `up` · `top` · `bottom` · `--to target` |
| `eval` | Run a JavaScript expression in the page and return its value | `javascript` (req.) |
| `screenshot` | Capture the viewport — or the **whole page** with `--full` — as **PNG** | `--file path`, `--full` |
| `html` | The `outerHTML` of an element (or the document) for markup debugging | `target` (optional), `--max N` |
| `pdf` | Save the page as **PDF** (headless only) | `--file path` (optional) |
| `resize` | Emulate a viewport (`--mobile` also enables touch) to verify responsive layouts | `width`, `height`, `--mobile` |
| `tabs` | List open tabs — popups such as OAuth windows show up here | — |
| `tab` | Switch to tab *n* from the `tabs` list | `n` (req.) |
| `cookies` | List cookies (**names and domains only, never values**) to check whether a login stuck; `--clear` logs out of everything | `domain` (optional), `--clear` |
| `console` | The captured console messages (errors and auto-accepted dialogs included) | `--tail N` |
| `network` | The captured network responses: method, status, URL | `--tail N` |
| `back` | History back | — |
| `status` | Whether a session is running, **visible or headless**, which profile it uses and what page it is on | — |
| `close` | Close the browser session | — |

***

## Refs vs. CSS selectors

Every `snapshot` stamps each interactive element with a `data-chatcli-ref` attribute and returns a **numbered listing**. `click` and `type` accept **either** that `[n]` ref number **or** a raw CSS selector:

```text theme={"system"}
Interactive elements:
  [1] link    "Docs"           → /docs
  [2] input   "Search"         (type=search)
  [3] button  "Sign in"

@browser click 3          # by ref (recommended — no guessing at selectors)
@browser click #login-btn # by CSS selector
```

<Note>
  Refs are stable **until the next navigation**. After you `open` a new URL or a `click` triggers navigation, take a fresh `snapshot` before addressing elements by `[n]` again.
</Note>

***

## Let the user log in (visible hand-off)

The agent's browser starts as a **throwaway profile**: none of the sessions in your everyday Chrome exist there. So when a task hits a login wall, the agent hands the window to you instead of guessing credentials:

```text theme={"system"}
# The agent surfaces the login page on your screen
@browser show https://app.example.com/login
  → Browser window is now VISIBLE on the user's screen. They can log in or
    interact directly; the session (cookies, logins) is shared with you…

# …and waits until the page reaches the post-login URL (up to 5 minutes)
@browser wait --url /dashboard --timeout 300
  → Condition met after 42s: Dashboard (https://app.example.com/dashboard)

# Window goes away; the authenticated session keeps working headless
@browser hide
@browser snapshot
```

`wait` accepts `--url`, `--text`, `--selector` and `--changed` (all must hold when combined). `--changed` fires when the URL leaves the one it had when the wait started — the right condition when the landing page cannot be predicted (GitHub, for instance, ignores the login's `return_to` and lands on its home page). A timeout is reported as a result, including where the page moved from if it did move, so the agent can ask you or wait again. Without a page condition to wait for, the agent asks you directly with [`@ask`](/coder/interactive-ask). Logins done this way live for the **rest of the ChatCLI session**; to keep them across runs, set `CHATCLI_BROWSER_PROFILE` (below).

<Tip>
  OAuth providers often open the sign-in flow in a **popup**. `tabs` lists it and `tab 2` lets the agent drive it — or you finish it yourself in the visible window.
</Tip>

<Note>
  **Signing in with Google.** Google refuses password sign-in in browsers it detects as automated ("This browser or app may not be secure"). ChatCLI launches Chrome with the automation marker off (`navigator.webdriver` is `false`), which is what that check looks at. If Google still refuses, sign into the site with a password or passkey instead of the Google button, or attach ChatCLI to the Chrome you are already signed into with `CHATCLI_BROWSER_CDP_URL` (below).
</Note>

If you **close the tab or the window** while the agent is waiting, `wait` reports it instead of failing, and the next `open`/`show` attaches a fresh tab — the browser session itself keeps running.

***

## A worked example

Build a page in `/coder`, then verify it end to end:

```text theme={"system"}
# 1) Open the app the coder just started
@browser open http://localhost:3000
  → Page: My App — Home
    Interactive elements:
      [1] link   "Products"
      [2] input  "Email"      (type=email)
      [3] button "Subscribe"

# 2) Fill the form and submit
@browser type 2 "user@example.com"
@browser click 3

# 3) See what happened
@browser snapshot
  → Page: My App — Thanks!
    "You're subscribed."

# 4) Something looked off? Debug from the page itself
@browser console --tail 20
  → [error] Uncaught TypeError: cannot read 'id' of undefined (app.js:42)

@browser network --tail 10
  → 500 POST http://localhost:3000/api/subscribe (Fetch)
```

The agent now knows the button worked in the UI but the backend returned **500**, and the exact console error — the full verify-and-fix loop without leaving ChatCLI.

***

## Environment variables

| Variable | Description | Default |
| :- | :- | :- |
| `CHATCLI_BROWSER_BIN` | Path to the browser binary to use | auto-detect Chrome/Chromium/Brave/Edge |
| `CHATCLI_BROWSER_HEADLESS` | Run headless (`true`) or open a **visible window** (`false`) for the whole session — `show`/`hide` switch per call regardless | `true` |
| `CHATCLI_BROWSER_PROFILE` | Where cookies and logins live. Unset = throwaway profile deleted on close. `true` = a persistent profile owned by ChatCLI under `~/.chatcli/browser/profile`; a path = that directory. Never your everyday Chrome profile | *(throwaway)* |
| `CHATCLI_BROWSER_CDP_URL` | **Attach to a browser you already have open** instead of launching one — `http://127.0.0.1:9222` (start Chrome with `--remote-debugging-port=9222`) or a `ws://` DevTools URL. ChatCLI opens one tab of its own there and closes only that tab on exit; the window is yours, so it is always visible and never killed | *(launch own browser)* |

```bash theme={"system"}
# Watch the agent drive a real window for the whole session
export CHATCLI_BROWSER_HEADLESS=false

# Keep logins between ChatCLI runs (profile owned by ChatCLI)
export CHATCLI_BROWSER_PROFILE=true

# Use the Chrome you are already signed into
# (launch it once with: google-chrome --remote-debugging-port=9222)
export CHATCLI_BROWSER_CDP_URL=http://127.0.0.1:9222

# Point at a specific browser
export CHATCLI_BROWSER_BIN="/Applications/Brave Browser.app/Contents/MacOS/Brave Browser"
```

***

## Security

Observation commands are **read-only** and skip the confirmation prompt — the agent can look freely: `open`, `show`, `hide`, `wait`, `snapshot`, `screenshot`, `html`, `pdf`, `resize`, `tabs`, `tab`, `cookies` (listing), `console`, `network`, `scroll`, `back`, `status`, `close`.

Commands that **act** on a page — `click`, `type`, `press`, `hover`, `select`, `upload`, `eval` and `cookies --clear` — can trigger real actions on remote sites (or log you out of everything), so they go through the standard **security gate** like any other mutating tool. See [Coder Mode Security](/coder/coder-security).

`cookies` never returns cookie **values**: the agent can tell that a session cookie exists for a domain, but a token never lands in the transcript.

<Warning>
  `@browser` drives a real browser signed into whatever profile it runs on. The default throwaway profile is the safe choice for untrusted pages; a persistent profile (`CHATCLI_BROWSER_PROFILE`) or your own browser (`CHATCLI_BROWSER_CDP_URL`) hands the agent every login stored there — treat `click`/`type`/`eval` on external sites as the side-effecting operations they are.
</Warning>

***

## Notes

* **One session per process**, launched lazily on the first `open` and reused across every call.
* Console and network are captured **continuously** into bounded rings — `console`/`network` read the recent history, they don't need to be "armed" first.
* `screenshot` writes a PNG (default under the temp dir, or `--file path`). Pair it with [`@view`](/usage/vision-input) so the model can **look at the screenshot** it just took.
* `alert()`, `confirm()` and `prompt()` dialogs are **accepted automatically** (they would otherwise freeze the page) and traced in `console` as `[dialog]` entries.
* A browser ChatCLI launched is **always** closed on ChatCLI shutdown; a stray headless Chrome never survives the session. An attached browser (`CHATCLI_BROWSER_CDP_URL`) only loses the tab ChatCLI opened.

## Next steps

<CardGroup cols={2}>
  <Card title="Image Input (@view)" icon="eye" href="/usage/vision-input">
    Attach a screenshot `@browser` captured so the model can see it with its own eyes.
  </Card>

  <Card title="Plugin @coder" icon="hammer" href="/coder/coder-plugin">
    Read, write, patch and run — pair it with `@browser` to build then verify.
  </Card>

  <Card title="Agentic Plugins" icon="plug" href="/agents/agentic-plugins">
    The full builtin tool catalog and how the agent uses them.
  </Card>

  <Card title="Web Tools" icon="globe" href="/tools/web-tools">
    `@webfetch`, `@websearch` and the shared hardened HTTP client.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.