optimize-agent-prompt

작성자: browserbase

Browserbase Agent API 데모를 구축하고 개선하며, Autobrowse 스타일의 외부 루프를 통해 고정된 작업을 실행하고, Agent 메시지와 세션 로그를 수집하고, 결과를 점수화하고, 하나의 시스템 프롬프트 휴리스틱을 수정하고, 수렴을 확인합니다. Browserbase Agents 데모 또는 POC를 만들 때, Agent 시스템 프롬프트를 최적화할 때, 불안정한 Agent 실행을 진단할 때, 또는 Browserbase Agents API에 자동 리서치/오토브라우즈를 적용할 때 사용하세요.

npx skills add https://github.com/browserbase/skills --skill optimize-agent-prompt

Optimize Agent Prompt

Optimize a Browserbase Agent's systemPrompt while holding its task, result schema, variables, and evaluation criteria fixed. Treat the outer agent as the teacher and each Browserbase Agent run as an inner-agent rollout.

Use Node.js 18 or later and set BROWSERBASE_API_KEY. The harness uses only Node.js built-in modules.

Set up the experiment

Choose a short experiment name and create an isolated workspace inside the demo or POC repository:

node <skill-dir>/scripts/optimize_agent_prompt.mjs init \
  --workspace ./agent-prompt-optimization/<experiment-name> \
  --name <experiment-name>

Edit the generated files:

  • task.json: keep task, resultSchema, variables, browser settings, and evaluation oracle stable across iterations.
  • prompts/iteration-001.md: write the minimal baseline system prompt. Include irreversible-action guardrails when applicable.

Use concrete success criteria. Prefer a strict JSON Schema with required fields and null for unavailable facts. Add known-field regexes and factuality-warning regexes under evaluation when a truth oracle exists. Read references/evaluation.md when designing the task or score.

Run the baseline

node <skill-dir>/scripts/optimize_agent_prompt.mjs run \
  --workspace ./agent-prompt-optimization/<experiment-name> \
  --prompt prompts/iteration-001.md \
  --label iteration-001

The harness creates one reusable Browserbase Agent, updates its systemPrompt on later iterations, starts the run, polls messages and status, and writes:

runs/<label>/
├── system-prompt.md
├── created-run.json
├── run.json
├── messages.json
├── session-logs.json
└── summary.json

It stops a run after the configured message budget instead of paying for an unproductive spiral. Use --max-messages, --timeout-ms, --proxies, or --verified only when the task needs different values from task.json.

Diagnose from observable evidence

Start with the compact trajectory:

node <skill-dir>/scripts/optimize_agent_prompt.mjs inspect \
  --workspace ./agent-prompt-optimization/<experiment-name> \
  --label iteration-001

Then read summary.json and drill into messages.json at the first wrong or wasted turn. Agent messages expose ordered tool calls, tool results, errors, and final output. A reasoning part may contain no readable text; never require hidden chain-of-thought for the teacher loop.

Read session-logs.json only when browser-level evidence can distinguish the cause—for example, a redirect, 403, failed request, console error, or hidden endpoint. Empty session logs can mean the Agent completed with search/fetch tools and never drove its browser.

See references/api.md for endpoint shapes, pagination, result normalization, and trace caveats.

Improve one heuristic

Find the earliest consequential failure and state one counterfactual:

If the system prompt had instructed X, the Agent would have avoided Y, as shown by tool result Z.

Copy the current prompt to prompts/iteration-NNN.md and make one attributable change. Typical improvements are:

  • cap retries after a repeated block or identical error;
  • distinguish public identifiers from private/internal IDs;
  • prefer search/fetch before launching a browser when interaction is unnecessary;
  • separate current snapshots from dated historical events;
  • define when a qualified fallback counts as completed;
  • require null instead of guessed values;
  • add a tool-call or evidence budget.

Keep wins. If the new run regresses, restore the previous prompt and test a different hypothesis rather than stacking more rules.

Judge and converge

Generate the comparison table after each run:

node <skill-dir>/scripts/optimize_agent_prompt.mjs report \
  --workspace ./agent-prompt-optimization/<experiment-name>

Judge more than field completeness. Require:

  • terminal status COMPLETED;
  • required fields populated or explicitly nullable;
  • known-fact checks passing when available;
  • no factuality-warning match;
  • provenance and safety constraints preserved;
  • fewer messages or lower duration without quality loss.

Once a prompt wins, run it again unchanged with a new label. Converge only after it passes at least two of the last three runs and one pass is an unchanged confirmation. Do not call a prompt globally optimal from one task; describe it as the best prompt for the tested task distribution.

Graduate into the demo

Use the confirmed prompt as the Agent's production systemPrompt. Keep the strict result schema and per-run variables. Preserve the experiment workspace or its report so reviewers can audit why each instruction exists.

In the final handoff, report:

  • baseline versus winning score, duration, and message count;
  • the first wrong turn each prompt change fixed;
  • whether session logs added evidence;
  • the winning prompt path;
  • confirmation-run results;
  • limitations and the next holdout matrix.

browserbase의 다른 스킬

add-webmcp
browserbase
기존 웹 애플리케이션을 분석하고, 라우트, 폼, 서버 액션, 핸들러, 스키마 전반에서 안전한 사용자 노출 기능을 식별한 다음, 자사 WebMCP 도구를 구현하고 Stagehand로 발견 및 호출을 검증합니다. 사용자가 코드베이스를 에이전트 준비 상태로 만들거나, 웹사이트 기능을 WebMCP 도구로 노출하거나, URL에서 독립 실행형 주입 스크립트를 생성하는 대신 앱에 WebMCP를 직접 추가하도록 요청할 때 사용합니다.
browse
browserbase
browse CLI를 사용하여 Browserbase 브라우저 자동화, Browserbase 클라우드 API, Browserbase Functions, 템플릿, 웹 가져오기/검색, 진단 및 Browse.sh…
browse
browserbase
browse CLI를 Browserbase 브라우저 자동화, Browserbase 클라우드 API, Browserbase Functions, 템플릿, 웹 fetch/search, 진단, Browse.sh 등에 사용하세요…
browser-automation
browserbase
MCP 도구를 사용하여 웹 브라우저 상호작용을 자동화합니다. 사용자가 웹사이트 탐색, 웹 페이지 이동, 웹사이트에서 데이터 추출, 스크린샷 촬영 등을 요청할 때 사용합니다,…
functions
browserbase
공식 Browserbase Functions CLI를 사용한 서버리스 브라우저 자동화 배포를 안내합니다. 사용자가 자동화를 배포하여 실행하려 할 때 사용하세요…
agent-experience
browserbase
제품, SDK, 문서 사이트 또는 SKILL.md의 개발자 경험을 감사하려면, 여러 Claude 하위 에이전트를 아주 작은 작업 프롬프트와 실제…
autobrowse
browserbase
자기 개선형 브라우저 자동화로, 자동 연구 루프를 통해 탐색 작업을 반복 실행하고, 추적 기록을 읽으며, 탐색 기술을 개선합니다…
browse
browserbase
browse CLI를 사용하여 브라우저 자동화 함수를 생성하고 배포하는 완벽 가이드