agent-dx-cli-scale

一個基於「為AI代理重寫你的CLI」原則,用於評估CLI對AI代理設計優劣的評分量表。

npx skills add https://github.com/google-labs-code/design.md --skill agent-dx-cli-scale

Agent DX CLI Scale

Use this skill to evaluate any CLI against the principles of agent-first design. Score each axis from 0–3, then sum for a total between 0–21.

Human DX optimizes for discoverability and forgiveness. Agent DX optimizes for predictability and defense-in-depth. — You Need to Rewrite Your CLI for AI Agents


Scoring Axes

1. Machine-Readable Output

Can an agent parse the CLI's output without heuristics?

ScoreCriteria
0Human-only output (tables, color codes, prose). No structured format available.
1--output json or equivalent exists but is incomplete or inconsistent across commands.
2Consistent JSON output across all commands. Errors also return structured JSON.
3NDJSON streaming for paginated results. Structured output is the default in non-TTY (piped) contexts.

2. Raw Payload Input

Can an agent send the full API payload without translation through bespoke flags?

ScoreCriteria
0Only bespoke flags. No way to pass structured input.
1Accepts --json or stdin JSON for some commands, but most require flags.
2All mutating commands accept a raw JSON payload that maps directly to the underlying API schema.
3Raw payload is first-class alongside convenience flags. The agent can use the API schema as documentation with zero translation loss.

3. Schema Introspection

Can an agent discover what the CLI accepts at runtime without pre-stuffed documentation?

ScoreCriteria
0Only --help text. No machine-readable schema.
1--help --json or a describe command for some surfaces, but incomplete.
2Full schema introspection for all commands — params, types, required fields — as JSON.
3Live, runtime-resolved schemas (e.g., from a discovery document) that always reflect the current API version. Includes scopes, enums, and nested types.

4. Context Window Discipline

Does the CLI help agents control response size to protect their context window?

ScoreCriteria
0Returns full API responses with no way to limit fields or paginate.
1Supports --fields or field masks on some commands.
2Field masks on all read commands. Pagination with --page-all or equivalent.
3Streaming pagination (NDJSON per page). Explicit guidance in context/skill files on field mask usage. The CLI actively protects the agent from token waste.

5. Input Hardening

Does the CLI defend against the specific ways agents fail (hallucinations, not typos)?

ScoreCriteria
0No input validation beyond basic type checks.
1Validates some inputs, but does not cover agent-specific hallucination patterns (path traversals, embedded query params, double encoding).
2Rejects control characters, path traversals (../), percent-encoded segments (%2e), and embedded query params (?, #) in resource IDs.
3Comprehensive hardening: all of the above, plus output path sandboxing to CWD, HTTP-layer percent-encoding, and an explicit security posture — "The agent is not a trusted operator."

6. Safety Rails

Can agents validate before acting, and are responses sanitized against prompt injection?

ScoreCriteria
0No dry-run mode. No response sanitization.
1--dry-run exists for some mutating commands.
2--dry-run for all mutating commands. Agent can validate requests without side effects.
3Dry-run plus response sanitization (e.g., via Model Armor) to defend against prompt injection embedded in API data. The full request→response loop is defended.

7. Agent Knowledge Packaging

Does the CLI ship knowledge in formats agents can consume at conversation start?

ScoreCriteria
0Only --help and a docs site. No agent-specific context files.
1A CONTEXT.md or AGENTS.md with basic usage guidance.
2Structured skill files (YAML frontmatter + Markdown) covering per-command or per-API-surface workflows and invariants.
3Comprehensive skill library encoding agent-specific guardrails ("always use --dry-run", "always use --fields"). Skills are versioned, discoverable, and follow a standard like OpenClaw.

Interpreting the Total

RangeRatingDescription
0–5Human-onlyBuilt for humans. Agents will struggle with parsing, hallucinate inputs, and lack safety rails.
6–10Agent-tolerantAgents can use it, but they'll waste tokens, make avoidable errors, and require heavy prompt engineering to compensate.
11–15Agent-readySolid agent support. Structured I/O, input validation, and some introspection. A few gaps remain.
16–21Agent-firstPurpose-built for agents. Full schema introspection, comprehensive input hardening, safety rails, and packaged agent knowledge.

Bonus: Multi-Surface Readiness

Not scored, but note whether the CLI exposes multiple agent surfaces from the same binary:

  • MCP (stdio JSON-RPC) — typed tool invocation, no shell escaping
  • Extension / plugin install — agent treats the CLI as a native capability
  • Headless auth — env vars for tokens/credentials, no browser redirect required

來自 google-labs-code 的更多技能

typed-service-contracts
google-labs-code
使用「Spec and Handler」模式建構穩健、型別安全的 TypeScript 服務之架構標準。適用於建置 CLI、函式庫或複雜…
stitch::code-to-design
google-labs-code
將前端程式碼(Vite、React 等)轉換為 Stitch Design,透過鏈式提取靜態 HTML、設計系統及檔案上傳。**務必**使用此…
react-vite-dashboard
google-labs-code
將Stitch設計轉換為生產級React + Vite儀表板,搭配TanStack Query、來自DESIGN.md的可存取token,以及Web3就緒模式(ethers/viem)。
stitch::extract-design-md
google-labs-code
從前端原始碼中直接提取全面的設計系統(DESIGN.md)——支援 React、Vue、Svelte、Angular、純 HTML/CSS 或任何網頁框架。
remotion
google-labs-code
使用 Remotion 從 Stitch 應用程式設計中建立專業的逐步解說影片,包含流暢的轉場與文字疊加。從 Stitch 專案中擷取畫面,並將其編排成 Remotion 影片組合,支援縮放效果、淡入淡出轉場及情境文字疊加。支援模組化元件架構,包含 ScreenSlide 與 WalkthroughComposition 元件,以及互動式熱點與旁白整合等進階功能。產生畫面清單、下載...
ink
google-labs-code
用於 json-render 的 Ink 終端渲染器,可將 JSON 規格轉換為互動式終端 UI。適用於使用 @json-render/ink 從…構建終端 UI 時。
tdd-red-green-refactor
google-labs-code
此技能實作了一個結構化框架,用於AI輔助程式設計,確保每一行程式碼都可驗證、具型別且有意義。
automate-github-issues
google-labs-code
使用並行的 Jules 編碼代理設置自動化的 GitHub 問題分類與解決。