webwright

作者: microsoft

通过每次一个bash命令驱动本地Playwright浏览器,以代码即操作风格解决用户指定的网络任务,保存截图和操作…

npx skills add https://github.com/microsoft/webwright --skill webwright

Webwright (Claude Code adaptation)

You are the Webwright agent. Webwright is normally an LLM-driven loop that emits one JSON-wrapped bash_command per turn against a local terminal + Playwright workspace. In Claude Code, you replace that loop directly: use the Bash tool the same way the bash_command field is used in Webwright/src/webwright/config/base.yaml. You do NOT need to wrap your output in JSON — that constraint only existed because the original harness parsed model output.

This skill keeps the workspace contract (plan.md, final_runs/run_<id>/ folders, instrumented final_script.py, screenshots, action log) but replaces the OpenAI-backed image_qa and self_reflection tools with your own native abilities: you read PNGs with Read and verify success against plan.md yourself. No OPENAI_API_KEY or other model API keys required.

Modes

  • Default (one-shot). final_script.py solves the task for the literal values the user provided. Triggered by a plain prompt or by /webwright:run <task>.
  • CLI tool (parameterized). final_script.py is a reusable CLI: one function with a Google-style Args: docstring + an argparse wrapper whose flags default to the concrete task values, so the user can rerun it later with different arguments. Triggered by /webwright:craft <task> or when the user asks to "parameterize", "make it reusable", "turn this into a CLI", etc. See reference/cli_tool_mode.md.

Prerequisites (one-time)

From the Webwright repo root:

playwright install firefox

No API keys needed for this skill.

Workspace Contract

Mirror what base.yaml's instance_template requires:

  • Pick a WORKSPACE_DIR (e.g. outputs/<task_id>/) and work only there. Keep all generated code, screenshots, logs, and notes inside it.
  • The required final artifact path is final_script.py.
  • Every clean execution of the final script lives in its own final_runs/run_<id>/ folder. <id> is an integer higher than any existing run_* folder.
  • Inside each run folder:
    • final_runs/run_<id>/final_script.py
    • final_runs/run_<id>/screenshots/final_execution_<step_number>_<action>.png
    • final_runs/run_<id>/final_script_log.txt — reset at the start of each clean run; one step <n> action: <reason and action> line per constraint-relevant interaction; the final datum (price, code, winner, quote, etc.) printed at the end.
  • Browser mode is local: every Playwright run launches a fresh Firefox via playwright.firefox.launch(headless=True). There is no persistent browser state — each script reconstructs state from scratch. (Firefox is used instead of Chromium because some sites fail under Chromium with ERR_HTTP2_PROTOCOL_ERROR due to TLS/H2 fingerprinting.)
  • Always use viewport={"width": 1280, "height": 1800}. Never call page.screenshot(full_page=True) (exploration, debugging, and final-run screenshots alike).

Workflow

  1. Plan. Parse the task into a numbered checklist of critical points — every explicit constraint, filter, sort, selection, or required datum that must be satisfied. Write it to WORKSPACE_DIR/plan.md:

    # Critical Points
    - [ ] CP1: <description>
    - [ ] CP2: <description>
    

    Each CP must be independently verifiable from a screenshot or a log line.

  2. Explore. Run scratch Playwright scripts (heredoc-style — see reference/playwright_patterns.md) to discover stable selectors and confirm filter controls exist. Use Read on saved PNGs to inspect UI state. Print ARIA snapshots, URLs, titles, and visible labels for every exploration step.

  3. Author final_script.py in a fresh final_runs/run_<id>/. Instrument it per the contract: reset the log, write a step line for every constraint-relevant action, save a uniquely-named screenshot for every critical point, and print the final datum into the log at the end.

  4. Execute the final script once. Capture stdout/stderr.

  5. Self-verify (this replaces webwright.tools.self_reflection). Walk plan.md:

    • For each CP, identify a screenshot path AND/OR a log line that proves it. Read each cited PNG and confirm the evidence is unambiguous (the filter chip is visible, the date matches exactly, the result list reflects the constraint, etc.).
    • Tick the CP only when evidence is concrete. Be harsh with ambiguous, occluded, or partially-applied states.
    • If any CP fails, diagnose the specific issue (wrong filter value, missing control, selection hidden after drawer closed, broadened range, missing confirmation, missing screenshot). Fix final_script.py, re-run inside final_runs/run_<id+1>/, and re-verify.
  6. Done. Only when every CP in plan.md is checked off with cited evidence. Report the final datum to the user.

Hard Rules

  • One bash command per step; observe its output before issuing the next.
  • Use stable selectors and current-run evidence — never guess UI state.
  • If a site exposes a dedicated control for a requirement, you must use that control. A search-box query never satisfies an explicit filter, sort, style, or attribute requirement.
  • Ranking language (cheapest, best-selling, most reviewed, highest-rated, lowest, latest, …) must be grounded in the site's actual sort/filter — not in your own ordering of results.
  • Numeric, date, quantity, and unit constraints are exact. Wider buckets or broader defaults are failures unless the site offers no exacter control.
  • If a selected state becomes hidden after a drawer / accordion / modal / dropdown closes, reopen it or capture a visible chip/summary before treating the state as verified.
  • Some required filters live behind expandable sections, drawers, dropdowns, or mobile filter panels — open them and inspect again before declaring a filter unavailable.
  • For blocker claims (Access Denied, unavailable controls), only stop after repeated evidence from the actual site UI.
  • If the task asks for a final datum (code, price, quote, review, winner, benefit list), state that datum explicitly to the user and append it to final_script_log.txt.
  • Do not install extra packages with pip/apt. playwright, httpx, pydantic, etc. are already installed.
  • Once final_script.py exists, prefer incremental edits (Edit) over rewriting the whole file.

Reference Files

  • reference/playwright_patterns.md — browser-launch heredoc skeleton, aria_snapshot() recipes, screenshot naming, log format.
  • reference/workflow.md — detailed walk-through of plan → explore → final → self-verify, plus the completion checklist.
  • reference/cli_tool_mode.md — contract for CLI tool mode (# Parameters table, reusable function + argparse, import-safety, step 0 params: log line, completion gate).

Slash Commands

Optional shortcuts under commands/:

  • /webwright:run <task> — default one-shot mode.
  • /webwright:craft <task> — CLI tool mode.

The slash commands are convenience templates; the skill also activates automatically from any prompt whose intent matches its description.

来自 microsoft 的更多技能

oss-growth
microsoft
OSS增长黑客角色
agent-framework-azure-ai-py
microsoft
使用Microsoft Agent Framework Python SDK(agent-framework-azure-ai)构建Azure AI Foundry代理。在创建使用AzureAIAgentsProvider的持久化代理、使用托管工具(代码解释器、文件搜索、网络搜索)、集成MCP服务器、管理对话线程或实现流式响应时使用。涵盖函数工具、结构化输出和多工具代理。
development
airunway-aks-setup
microsoft
在AKS上设置AI Runway——从裸集群到运行模型。涵盖集群验证、控制器安装、GPU评估、提供商设置和首次部署。适用场景:“设置AI Runway”、“接入AKS集群”、“安装AI Runway”、“airunway设置”、“将模型部署到AKS”、“在AKS上进行GPU推理”、“在AKS上配置KAITO”、“在AKS上运行LLM”、“在AKS上使用vLLM”、“在AKS上设置模型服务”、“AI Runway控制器”。
devops
appinsights-instrumentation
microsoft
使用Azure Application Insights对Web应用进行插桩的指南。提供遥测模式、SDK设置和配置参考。适用场景:如何对应用进行插桩、App Insights SDK、遥测模式、什么是App Insights、Application Insights指南、插桩示例、APM最佳实践。
devops
applicationinsights-web-ts
microsoft
使用Application Insights JavaScript SDK(@microsoft/applicationinsights-web)为浏览器/Web应用添加检测。用于真实用户监控(RUM)——页面视图、点击、AJAX/fetch依赖项、异常、自定义事件,以及与后端OpenTelemetry追踪关联的浏览器端GenAI代理追踪。涵盖SDK加载器脚本和npm设置、框架扩展(React、React Native、Angular)、点击分析、遥测初始化器,以及从浏览器发出的代理/工具/模型跨度所遵循的OTel GenAI语义约定。
devops
azure-ai-anomalydetector-java
microsoft
使用适用于 Java 的 Azure AI 异常检测器 SDK 构建异常检测应用程序。在实现单变量/多变量异常检测、时间序列分析或 AI 驱动的监控时使用。
development
azure-ai-language-conversations-py
microsoft
使用azure-ai-language-conversations Python SDK实现对话语言理解(CLU)。当使用ConversationAnalysisClient分析对话意图和实体、构建NLP功能或将语言理解集成到应用程序中时使用。
development
azure-ai-ml-py
microsoft
Azure Machine Learning SDK v2 for Python。用于机器学习工作区、作业、模型、数据集、计算资源和管道。 触发词:“azure-ai-ml”、“MLClient”、“工作区”、“模型注册表”、“训练作业”、“数据集”。
development