browser

作者: browserbase

使用本地Chrome或远程Browserbase进行浏览器自动化,适用于受保护网站、机器人检测和验证码场景。两种模式:本地Chrome(默认,无需配置)或远程Browserbase(反机器人隐身、自动验证码破解、住宅代理、会话持久化)。核心命令涵盖导航、页面检查、交互(点击、输入、填充、选择、拖拽)以及通过CLI进行会话管理。使用browse snapshot读取无障碍树并获取元素引用以实现可靠交互;保留...

npx skills add https://github.com/browserbase/skills --skill browser

Browser Automation

Automate browser interactions using the browse CLI with Claude.

Setup check

Before running any browser commands, verify the CLI is available:

which browse || npm install -g browse

Environment Selection (Local vs Remote)

The CLI supports explicit per-command environment flags. If you do nothing, the next session defaults to Browserbase when BROWSERBASE_API_KEY is set and to local otherwise.

Local mode

  • browse open <url> --local starts a clean isolated local browser
  • browse open <url> --auto-connect attaches to an already-running debuggable Chrome; use --local when no debuggable Chrome is available
  • browse open <url> --cdp <port|url> attaches to a specific CDP target
  • Best for: development, localhost, trusted sites, and reproducible runs

Remote mode (Browserbase)

  • browse open <url> --remote starts a Browserbase session
  • Without a local flag, Browserbase is also the default when BROWSERBASE_API_KEY is set
  • Provides: Browserbase Identity, Verified browsers, automatic CAPTCHA solving, residential proxies, session persistence
  • Use remote mode when: the target site has bot detection, CAPTCHAs, IP rate limiting, Cloudflare protection, or requires geo-specific access
  • Get credentials at https://browserbase.com/settings

When to choose which

  • Repeatable local testing / clean state: browse open <url> --local
  • Reuse your local login/cookies: browse open <url> --auto-connect
  • Simple browsing (docs, wikis, public APIs): local mode is fine
  • Protected sites (login walls, CAPTCHAs, anti-scraping): use remote mode
  • If local mode fails with bot detection or access denied: switch to remote mode

Commands

Most driver commands work across local, remote, and CDP sessions after the daemon starts.

Navigation

browse open <url>                        # Go to URL
browse open <url> --local                # Go to URL in a clean local browser
browse open <url> --remote               # Go to URL in a Browserbase session
browse reload                            # Reload current page
browse back                              # Go back in history
browse forward                           # Go forward in history

Page state (prefer snapshot over screenshot)

browse snapshot                          # Get accessibility tree with element refs (fast, structured)
browse screenshot --path <path>          # Take visual screenshot (slow, uses vision tokens)
browse get url                           # Get current URL
browse get title                         # Get page title
browse get text <selector>               # Get text content (use "body" for all text)
browse get html <selector>               # Get HTML content of element
browse get markdown [selector]           # Get page content as markdown (defaults to body)
browse get value <selector>              # Get form field value

Use browse snapshot as your default for understanding page state — it returns the accessibility tree with element refs you can use to interact. Only use browse screenshot when you need visual context (layout, images, debugging).

Interaction

browse click <ref>                       # Click element by ref from snapshot (e.g., @0-5)
browse type <text>                       # Type text into focused element
browse fill <selector> <value>           # Fill input; add --press-enter if Enter is needed
browse select <selector> <values...>     # Select dropdown option(s)
browse upload <selector> <files...>      # Upload file(s) to <input type="file">
browse press <key>                       # Press key (Enter, Tab, Escape, Cmd+A, etc.)
browse mouse drag <fromX> <fromY> <toX> <toY>  # Drag from one point to another
browse mouse scroll <x> <y> <deltaX> <deltaY>  # Scroll at coordinates
browse highlight <selector>              # Highlight element on page
browse is visible <selector>             # Check if element is visible
browse is checked <selector>             # Check if element is checked
browse wait <type> [arg]                 # Wait for: load, selector, timeout

CDP event tailing

browse cdp <url|port>                    # Stream CDP events as NDJSON from any target
browse cdp 9222                          # Attach to local Chrome on port 9222
browse cdp ws://localhost:9222/devtools/browser/...  # Full WebSocket URL
browse cdp <url> --domain Network        # Only Network events
browse cdp <url> --domain Network --domain Console  # Multiple domains
browse cdp <url> --pretty                # Human-readable output
browse cdp <url> > events.jsonl          # Pipe to file
browse cdp <url> | jq '.method'          # Filter with jq

The cdp command connects directly to any Chrome DevTools Protocol target and streams events. It does not use the daemon — it's a standalone, long-running process. Press Ctrl+C to stop. Default domains: Network, Console, Runtime, Log, Page.

Session management

browse stop                              # Stop the browser daemon
browse status                            # Check daemon status and resolved mode
browse tab list                          # List all open tabs
browse tab switch <index-or-target-id>   # Switch to tab by index or target ID
browse tab close [index-or-target-id]    # Close tab

Typical workflow

If the environment matters, put --local, --remote, --auto-connect, or --cdp <port|url> on the first browser command.

  1. browse open <url> --local or browse open <url> --remote — navigate to the page
  2. browse snapshot — read the accessibility tree to understand page structure and get element refs
  3. browse click <ref> / browse type <text> / browse fill <selector> <value> — interact using refs from snapshot
  4. browse snapshot — confirm the action worked
  5. Repeat 3-4 as needed
  6. browse stop — close the browser when done

Quick Example

browse open https://example.com
browse snapshot                          # see page structure + element refs
browse click @0-5                        # click element with ref 0-5
browse get title
browse stop

Mode Comparison

FeatureLocalBrowserbase
SpeedFasterSlightly slower
SetupChrome requiredAPI key required
Reuse existing local cookiesWith browse open <url> --auto-connectN/A
Verified browserNoYes (Browserbase Verified browser via Identity)
CAPTCHA solvingNoYes (automatic reCAPTCHA/hCaptcha)
Residential proxiesNoYes (201 countries, geo-targeting)
Session persistenceNoYes (cookies/auth persist via contexts)
Best forDevelopment/simple pagesProtected sites, Browserbase Identity + Verified access, production scraping

Best Practices

  1. Choose the local strategy deliberately: use browse open <url> --local for clean state, browse open <url> --auto-connect for existing local credentials, and browse open <url> --remote for protected sites
  2. Always browse open first before interacting
  3. Use browse snapshot to check page state — it's fast and gives you element refs
  4. Only screenshot when visual context is needed (layout checks, images, debugging)
  5. Use refs from snapshot to click/interact — e.g., browse click @0-5
  6. browse stop when done to clean up the browser session and clear the env override

Troubleshooting

  • "No active page": Run browse stop, then check browse status. If it still says running, kill the zombie daemon with pkill -f "browse.*daemon", then retry browse open
  • Chrome not found: Install Chrome, use browse open <url> --auto-connect if you already have a debuggable Chrome running, or switch to browse open <url> --remote
  • Action fails: Run browse snapshot to see available elements and their refs
  • Browserbase fails: Verify API key is set

Switching to Remote Mode

Switch to remote when you detect: CAPTCHAs (reCAPTCHA, hCaptcha, Turnstile), bot detection pages ("Checking your browser..."), HTTP 403/429, empty pages on sites that should have content, or the user asks for it.

Don't switch for simple sites (docs, wikis, public APIs, localhost).

browse open <url> --local          # clean isolated local browser
browse open <url> --auto-connect   # attach to existing debuggable Chrome
browse open <url> --remote         # Browserbase session

Mode flags are applied when a session starts. After browse stop, the next start falls back to env-var-based auto detection. Use browse status to inspect the resolved mode and target while the daemon is running.

For detailed examples, see EXAMPLES.md. For API reference, see REFERENCE.md.

来自 browserbase 的更多技能

optimize-agent-prompt
browserbase
通过Autobrowse式外层循环构建并改进Browserbase Agent API演示:运行固定任务,收集Agent消息和会话日志,对结果评分,修订一条系统提示启发式规则,并确认收敛。适用于创建Browserbase Agents演示或概念验证、优化Agent系统提示、诊断Agent运行不稳定,或将自动研究/自动浏览应用于Browserbase Agents API时。
add-webmcp
browserbase
分析现有Web应用,识别路由、表单、服务器操作、处理器和模式中安全且用户可见的功能,然后实现第一方WebMCP工具,并使用Stagehand验证发现和调用。当用户要求使代码库具备代理就绪性、将网站功能暴露为WebMCP工具,或直接将WebMCP添加到应用而非从URL生成独立注入脚本时使用。
browse
browserbase
使用browse CLI进行Browserbase浏览器自动化、Browserbase云API、Browserbase Functions、模板、网页抓取/搜索、诊断以及Browse.sh…
browse
browserbase
使用 browse CLI 进行 Browserbase 浏览器自动化、Browserbase 云 API、Browserbase Functions、模板、网页抓取/搜索、诊断以及 Browse.sh…
browser-automation
browserbase
使用MCP工具自动化网页浏览器交互。当用户要求浏览网站、导航网页、从网站提取数据、截图时使用,……
functions
browserbase
引导使用官方 Browserbase Functions CLI 进行无服务器浏览器自动化的部署。当用户想要部署自动化以在……上运行时使用。
agent-experience
browserbase
通过仅使用一个简短的任务提示和真实的……,投放多个Claude子代理来审计产品、SDK、文档站点或SKILL.md的开发者体验。
autobrowse
browserbase
通过自动研究循环实现自我改进的浏览器自动化。迭代执行浏览任务、读取追踪记录并优化导航技能…