optimize-agent-prompt

作者: browserbase

透過Autobrowse風格的外部迴圈來建置並改進Browserbase Agent API示範:執行固定任務、收集Agent訊息與工作階段日誌、評分結果、修訂一條系統提示啟發式規則,並確認收斂。適用於建立Browserbase Agents示範或概念驗證、最佳化Agent系統提示、診斷不穩定的Agent執行,或將自動研究/自動瀏覽應用於Browserbase Agents API時。

npx skills add https://github.com/browserbase/skills --skill optimize-agent-prompt

Optimize Agent Prompt

Optimize a Browserbase Agent's systemPrompt while holding its task, result schema, variables, and evaluation criteria fixed. Treat the outer agent as the teacher and each Browserbase Agent run as an inner-agent rollout.

Use Node.js 18 or later and set BROWSERBASE_API_KEY. The harness uses only Node.js built-in modules.

Set up the experiment

Choose a short experiment name and create an isolated workspace inside the demo or POC repository:

node <skill-dir>/scripts/optimize_agent_prompt.mjs init \
  --workspace ./agent-prompt-optimization/<experiment-name> \
  --name <experiment-name>

Edit the generated files:

  • task.json: keep task, resultSchema, variables, browser settings, and evaluation oracle stable across iterations.
  • prompts/iteration-001.md: write the minimal baseline system prompt. Include irreversible-action guardrails when applicable.

Use concrete success criteria. Prefer a strict JSON Schema with required fields and null for unavailable facts. Add known-field regexes and factuality-warning regexes under evaluation when a truth oracle exists. Read references/evaluation.md when designing the task or score.

Run the baseline

node <skill-dir>/scripts/optimize_agent_prompt.mjs run \
  --workspace ./agent-prompt-optimization/<experiment-name> \
  --prompt prompts/iteration-001.md \
  --label iteration-001

The harness creates one reusable Browserbase Agent, updates its systemPrompt on later iterations, starts the run, polls messages and status, and writes:

runs/<label>/
├── system-prompt.md
├── created-run.json
├── run.json
├── messages.json
├── session-logs.json
└── summary.json

It stops a run after the configured message budget instead of paying for an unproductive spiral. Use --max-messages, --timeout-ms, --proxies, or --verified only when the task needs different values from task.json.

Diagnose from observable evidence

Start with the compact trajectory:

node <skill-dir>/scripts/optimize_agent_prompt.mjs inspect \
  --workspace ./agent-prompt-optimization/<experiment-name> \
  --label iteration-001

Then read summary.json and drill into messages.json at the first wrong or wasted turn. Agent messages expose ordered tool calls, tool results, errors, and final output. A reasoning part may contain no readable text; never require hidden chain-of-thought for the teacher loop.

Read session-logs.json only when browser-level evidence can distinguish the cause—for example, a redirect, 403, failed request, console error, or hidden endpoint. Empty session logs can mean the Agent completed with search/fetch tools and never drove its browser.

See references/api.md for endpoint shapes, pagination, result normalization, and trace caveats.

Improve one heuristic

Find the earliest consequential failure and state one counterfactual:

If the system prompt had instructed X, the Agent would have avoided Y, as shown by tool result Z.

Copy the current prompt to prompts/iteration-NNN.md and make one attributable change. Typical improvements are:

  • cap retries after a repeated block or identical error;
  • distinguish public identifiers from private/internal IDs;
  • prefer search/fetch before launching a browser when interaction is unnecessary;
  • separate current snapshots from dated historical events;
  • define when a qualified fallback counts as completed;
  • require null instead of guessed values;
  • add a tool-call or evidence budget.

Keep wins. If the new run regresses, restore the previous prompt and test a different hypothesis rather than stacking more rules.

Judge and converge

Generate the comparison table after each run:

node <skill-dir>/scripts/optimize_agent_prompt.mjs report \
  --workspace ./agent-prompt-optimization/<experiment-name>

Judge more than field completeness. Require:

  • terminal status COMPLETED;
  • required fields populated or explicitly nullable;
  • known-fact checks passing when available;
  • no factuality-warning match;
  • provenance and safety constraints preserved;
  • fewer messages or lower duration without quality loss.

Once a prompt wins, run it again unchanged with a new label. Converge only after it passes at least two of the last three runs and one pass is an unchanged confirmation. Do not call a prompt globally optimal from one task; describe it as the best prompt for the tested task distribution.

Graduate into the demo

Use the confirmed prompt as the Agent's production systemPrompt. Keep the strict result schema and per-run variables. Preserve the experiment workspace or its report so reviewers can audit why each instruction exists.

In the final handoff, report:

  • baseline versus winning score, duration, and message count;
  • the first wrong turn each prompt change fixed;
  • whether session logs added evidence;
  • the winning prompt path;
  • confirmation-run results;
  • limitations and the next holdout matrix.

來自 browserbase 的更多技能

add-webmcp
browserbase
分析現有的網頁應用程式,識別跨路由、表單、伺服器動作、處理器和架構中的安全使用者可見功能,然後實作第一方 WebMCP 工具,並使用 Stagehand 驗證探索和呼叫。當使用者要求讓程式碼庫具備代理程式就緒狀態、將網站功能公開為 WebMCP 工具,或直接將 WebMCP 新增至應用程式(而非從 URL 產生獨立的注入腳本)時使用。
browse
browserbase
使用 browse CLI 進行 Browserbase 瀏覽器自動化、Browserbase 雲端 API、Browserbase Functions、範本、網頁擷取/搜尋、診斷,以及 Browse.sh…
browse
browserbase
使用 browse CLI 進行 Browserbase 瀏覽器自動化、Browserbase 雲端 API、Browserbase Functions、模板、網頁擷取/搜尋、診斷,以及 Browse.sh…
browser-automation
browserbase
使用MCP工具自動化網頁瀏覽器操作。當使用者要求瀏覽網站、導覽網頁、從網站提取資料、擷取螢幕截圖等時使用。
functions
browserbase
使用官方 Browserbase Functions CLI 引導無伺服器瀏覽器自動化的部署。當使用者想要部署自動化以在…上執行時使用。
agent-experience
browserbase
審核產品、SDK、文檔網站或 SKILL.md 的開發者體驗,方法是投放多個 Claude 子代理,僅附上極簡任務提示和真實…
autobrowse
browserbase
透過自動研究循環實現自我改進的瀏覽器自動化。反覆執行瀏覽任務、讀取追蹤記錄,並持續提升導航技能…
browse
browserbase
使用 browse CLI 建立和部署瀏覽器自動化功能的完整指南