Skip to main content
The @browser tool drives a real local Chrome/Chromium browser so the agent can see and interact with web pages the way a person would: open a URL, read the rendered page as text, click and type into elements, run JavaScript, capture screenshots, and inspect the page’s console and network activity. It is the verification loop for web work — the agent builds or changes a frontend, then actually looks at it and debugs it from what the page logged and requested, instead of guessing from source.
@browser speaks the Chrome DevTools Protocol (CDP) directly over the websocket client ChatCLI already ships — zero new dependency, no driver, no downloaded runtime, no API key. It only needs a locally installed Chromium-family browser (Google Chrome, Chromium, Brave or Edge).

How it works

  1. Lazy launch — the first open starts one browser; every later command reuses that single stateful session.
  2. CDP over websocket — navigation, DOM reads, screenshots and event capture all flow over the DevTools websocket. No Selenium/Playwright, no external binary.
  3. Snapshot as text — the page is rendered to model-friendly text with every interactive element stamped and numbered so the agent can address them precisely.
  4. Continuous capture — console messages (errors included) and network responses are captured into bounded rings, so you can ask “what did the page log?” after the fact.
  5. Visible on demand — the browser is headless by default, but the agent can put the window on your screen with show (or open --visible) whenever you need to act — log in, pick an account, solve a captcha — then take it back with hide. The switch relaunches the browser on the same profile, so cookies and logins survive it.
  6. Dies with the session — a browser ChatCLI launched never outlives it; it is closed on CLI shutdown (or explicitly with close).

Subcommands

Invoke with a JSON envelope {cmd, args} or the flat argv form. The model calls @browser automatically in /agent and /coder when it needs to look at a page.

Refs vs. CSS selectors

Every snapshot stamps each interactive element with a data-chatcli-ref attribute and returns a numbered listing. click and type accept either that [n] ref number or a raw CSS selector:
Refs are stable until the next navigation. After you open a new URL or a click triggers navigation, take a fresh snapshot before addressing elements by [n] again.

Let the user log in (visible hand-off)

The agent’s browser starts as a throwaway profile: none of the sessions in your everyday Chrome exist there. So when a task hits a login wall, the agent hands the window to you instead of guessing credentials:
wait accepts --url, --text, --selector and --changed (all must hold when combined). --changed fires when the URL leaves the one it had when the wait started — the right condition when the landing page cannot be predicted (GitHub, for instance, ignores the login’s return_to and lands on its home page). A timeout is reported as a result, including where the page moved from if it did move, so the agent can ask you or wait again. Without a page condition to wait for, the agent asks you directly with @ask. Logins done this way live for the rest of the ChatCLI session; to keep them across runs, set CHATCLI_BROWSER_PROFILE (below).
OAuth providers often open the sign-in flow in a popup. tabs lists it and tab 2 lets the agent drive it — or you finish it yourself in the visible window.
Signing in with Google. Google refuses password sign-in in browsers it detects as automated (“This browser or app may not be secure”). ChatCLI launches Chrome with the automation marker off (navigator.webdriver is false), which is what that check looks at. If Google still refuses, sign into the site with a password or passkey instead of the Google button, or attach ChatCLI to the Chrome you are already signed into with CHATCLI_BROWSER_CDP_URL (below).
If you close the tab or the window while the agent is waiting, wait reports it instead of failing, and the next open/show attaches a fresh tab — the browser session itself keeps running.

A worked example

Build a page in /coder, then verify it end to end:
The agent now knows the button worked in the UI but the backend returned 500, and the exact console error — the full verify-and-fix loop without leaving ChatCLI.

Environment variables


Security

Observation commands are read-only and skip the confirmation prompt — the agent can look freely: open, show, hide, wait, snapshot, screenshot, html, pdf, resize, tabs, tab, cookies (listing), console, network, scroll, back, status, close. Commands that act on a page — click, type, press, hover, select, upload, eval and cookies --clear — can trigger real actions on remote sites (or log you out of everything), so they go through the standard security gate like any other mutating tool. See Coder Mode Security. cookies never returns cookie values: the agent can tell that a session cookie exists for a domain, but a token never lands in the transcript.
@browser drives a real browser signed into whatever profile it runs on. The default throwaway profile is the safe choice for untrusted pages; a persistent profile (CHATCLI_BROWSER_PROFILE) or your own browser (CHATCLI_BROWSER_CDP_URL) hands the agent every login stored there — treat click/type/eval on external sites as the side-effecting operations they are.

Notes

  • One session per process, launched lazily on the first open and reused across every call.
  • Console and network are captured continuously into bounded rings — console/network read the recent history, they don’t need to be “armed” first.
  • screenshot writes a PNG (default under the temp dir, or --file path). Pair it with @view so the model can look at the screenshot it just took.
  • alert(), confirm() and prompt() dialogs are accepted automatically (they would otherwise freeze the page) and traced in console as [dialog] entries.
  • A browser ChatCLI launched is always closed on ChatCLI shutdown; a stray headless Chrome never survives the session. An attached browser (CHATCLI_BROWSER_CDP_URL) only loses the tab ChatCLI opened.

Next steps

Image Input (@view)

Attach a screenshot @browser captured so the model can see it with its own eyes.

Plugin @coder

Read, write, patch and run — pair it with @browser to build then verify.

Agentic Plugins

The full builtin tool catalog and how the agent uses them.

Web Tools

@webfetch, @websearch and the shared hardened HTTP client.