improving-mcp-tools

作成者: posthog

improve-my-MCPキャンペーンを実行する:評価ハーネスでMCPエージェント体験を測定し、最も影響の大きいツール問題を選定する自動研究スタイルのループ…

npx skills add https://github.com/posthog/ai-plugin --skill improving-mcp-tools

Improving MCP tools

An MCP server gets better only in ways you can measure. This skill is the campaign procedure: score the current agent experience, fix the biggest problem, re-score, and only ship changes the numbers justify. It is the operating manual for the "improve my MCP" loop — one iteration per pass, journaled so a later iteration (or a different agent) can resume without repeating work.

The objective function

services/mcp/evals/ is the harness. benchmark/tasks.yaml is a fixed set of agent tasks with expected_tools and success_criteria; scores are only comparable across runs of the same benchmark version.

  • Probe mode (deterministic, no LLM): LIVE_MCP_URL=... LIVE_MCP_TOKEN=... pnpm exec tsx evals/runner/probe.ts --out score.json from services/mcp/. Reports tool-presence misses (discoverability), probe failures, and latency p50/p95. Non-zero exit = regression.
  • Agent mode (LLM replay + judge): scores task success and tool-selection accuracy. Use it for description/discoverability changes — probes cannot detect that an agent picks the wrong tool.

Run the harness against a seeded local or devbox stack, never against a customer project. Local recipe: NODE_ENV=development PORT=9876 POSTHOG_API_BASE_URL=http://localhost:8000 pnpm dev:hono, personal API key as LIVE_MCP_TOKEN.

One iteration

  1. Measure. Run the harness for a baseline. Pull production evidence with the MCP analytics tools (query-mcp-tool-stats, query-mcp-tool-failures, query-mcp-tool-descriptions, query-mcp-tool-sample-intents) and the lenses in the signals scout cookbook (products/signals/skills/signals-scout-mcp-tool-calls/references/queries.md): failure leaderboard, retry/struggle, latency, intents that matched no tool.
  2. Pick one issue. Rank by reach × severity. Skip anything the journal shows with two failed attempts. One issue per iteration — a PR that fixes three things can't be attributed to any of them when scores move.
  3. Fix, bounded. Only files inside the allowlist (below). Typical fixes: sharpen a tool description so the right intent finds it, tighten an input schema that agents keep getting wrong, fix an annotation, update a skill.
  4. Validate. Re-run the affected benchmark slice plus a no-regression sample. Keep the change only if the target metric improves and nothing else degrades. A discarded change is a normal outcome — journal it and move on.
  5. Ship. One PR per iteration with before/after scores in the body (format in references/campaign-journal.md). Keep it stampable: ≤400 changed lines, only files inside the allowlist below, apply the stamphog label. Autonomy level comes from the campaign config — default is draft PR for human review; only arm auto-merge when the operator has explicitly enabled the self-driving experiment (see guardrails).
  6. Journal. Append the iteration record before ending the pass.

Hard guardrails

These are not suggestions; violating any of them ends the campaign pass.

  • Allowlist — a campaign PR may only touch: products/*/mcp/tools.yaml, products/*/skills/**, services/mcp/evals/**, the codegen outputs of pnpm generate-tools / scaffold-yaml (services/mcp/src/tools/generated/** and services/mcp/schema/generated-tool-definitions.json), and docs. Anything else (handler code, package manifests, workflows, migrations, auth paths) → stop and hand the finding to a human as a draft PR or report instead.
  • Read-only against data. The harness and all production queries are read-only. Never create, mutate, or delete customer-visible objects while measuring.
  • Evidence or it didn't happen. No PR without a baseline score, an after score, and the exact harness commands used.
  • Benchmark integrity. Never edit benchmark/tasks.yaml in the same PR as a fix it validates — changing the exam and the answer together proves nothing. Benchmark changes are their own PR and bump version.
  • Budgets. Respect the operator's iteration/token/PR caps (default: stop after 3 open unmerged campaign PRs). Two failed attempts on an issue parks it permanently.
  • Kill switch. If the campaign config, its feature flag, or the operator says stop — stop mid-iteration, journal state, end cleanly.

Failure modes to expect

  • A description change that helps one intent can steal traffic from the right tool for another — that's why the no-regression sample is mandatory. The intent-cluster snapshot's tool_overlaps (see exploring-mcp-intent-clusters) lists exactly which pairs compete for which intents: snapshot it before a description rewrite and recompute after, and treat a capture shift in an overlapping pair as the regression signal.
  • Probe latency varies with stack warmth; compare medians across ≥3 runs before attributing a latency change to your fix.
  • Tool-presence misses can be feature-flag gating, not catalog absence — check getToolsForFeatures gating before "fixing" discoverability.

posthogのその他のスキル

managing-experiment-lifecycle
posthog
実験の状態遷移をガイドします:起動、一時停止、再開、終了、バリアントの出荷、アーカイブ、リセット、複製。前提条件などをカバーします。
official
configuring-experiment-analytics
posthog
Configures the analytics side of a PostHog experiment — exposure criteria (default `$feature_flag_called` vs custom exposure events), primary and secondary…
official
error-tracking-hono
posthog
PostHogのHono向けエラートラッキング
official
error-tracking-react
posthog
PostHogのReact向けエラートラッキング
official
integration-android
posthog
Androidアプリケーション向けPostHogインテグレーション
official
integration-ruby
posthog
PostHog統合:Ruby SDKを使用するRubyアプリケーション向け
official
tuning-incremental-sync-config
posthog
同期の設定はExternalDataSchemaに保存され、external-data-schemas-partial-updateを使用していつでも変更できます。ほとんどの変更は非破壊的(次の同期で反映)ですが、一部(sync_typeの切り替え、プライマリキーの変更)は、同期データの破損を防ぐために慎重な対応が必要です。
official
instrument-integration
posthog
このスキルを使用して、アプリケーションにPostHog SDKを追加します。PostHogを初めてセットアップする場合や、PostHogの初期化が必要なPRをレビューする場合に使用してください。SDKのインストール、プロバイダーのセットアップ、基本設定をカバーします。任意のフレームワークや言語に対応しています。
official