agent-dx-cli-scale

작성자: google-labs-code

CLI가 AI 에이전트에 얼마나 잘 설계되었는지 평가하기 위한 점수 척도로, "AI 에이전트를 위한 CLI 재작성" 원칙에 기반합니다.

npx skills add https://github.com/google-labs-code/design.md --skill agent-dx-cli-scale

Agent DX CLI Scale

Use this skill to evaluate any CLI against the principles of agent-first design. Score each axis from 0–3, then sum for a total between 0–21.

Human DX optimizes for discoverability and forgiveness. Agent DX optimizes for predictability and defense-in-depth. — You Need to Rewrite Your CLI for AI Agents


Scoring Axes

1. Machine-Readable Output

Can an agent parse the CLI's output without heuristics?

ScoreCriteria
0Human-only output (tables, color codes, prose). No structured format available.
1--output json or equivalent exists but is incomplete or inconsistent across commands.
2Consistent JSON output across all commands. Errors also return structured JSON.
3NDJSON streaming for paginated results. Structured output is the default in non-TTY (piped) contexts.

2. Raw Payload Input

Can an agent send the full API payload without translation through bespoke flags?

ScoreCriteria
0Only bespoke flags. No way to pass structured input.
1Accepts --json or stdin JSON for some commands, but most require flags.
2All mutating commands accept a raw JSON payload that maps directly to the underlying API schema.
3Raw payload is first-class alongside convenience flags. The agent can use the API schema as documentation with zero translation loss.

3. Schema Introspection

Can an agent discover what the CLI accepts at runtime without pre-stuffed documentation?

ScoreCriteria
0Only --help text. No machine-readable schema.
1--help --json or a describe command for some surfaces, but incomplete.
2Full schema introspection for all commands — params, types, required fields — as JSON.
3Live, runtime-resolved schemas (e.g., from a discovery document) that always reflect the current API version. Includes scopes, enums, and nested types.

4. Context Window Discipline

Does the CLI help agents control response size to protect their context window?

ScoreCriteria
0Returns full API responses with no way to limit fields or paginate.
1Supports --fields or field masks on some commands.
2Field masks on all read commands. Pagination with --page-all or equivalent.
3Streaming pagination (NDJSON per page). Explicit guidance in context/skill files on field mask usage. The CLI actively protects the agent from token waste.

5. Input Hardening

Does the CLI defend against the specific ways agents fail (hallucinations, not typos)?

ScoreCriteria
0No input validation beyond basic type checks.
1Validates some inputs, but does not cover agent-specific hallucination patterns (path traversals, embedded query params, double encoding).
2Rejects control characters, path traversals (../), percent-encoded segments (%2e), and embedded query params (?, #) in resource IDs.
3Comprehensive hardening: all of the above, plus output path sandboxing to CWD, HTTP-layer percent-encoding, and an explicit security posture — "The agent is not a trusted operator."

6. Safety Rails

Can agents validate before acting, and are responses sanitized against prompt injection?

ScoreCriteria
0No dry-run mode. No response sanitization.
1--dry-run exists for some mutating commands.
2--dry-run for all mutating commands. Agent can validate requests without side effects.
3Dry-run plus response sanitization (e.g., via Model Armor) to defend against prompt injection embedded in API data. The full request→response loop is defended.

7. Agent Knowledge Packaging

Does the CLI ship knowledge in formats agents can consume at conversation start?

ScoreCriteria
0Only --help and a docs site. No agent-specific context files.
1A CONTEXT.md or AGENTS.md with basic usage guidance.
2Structured skill files (YAML frontmatter + Markdown) covering per-command or per-API-surface workflows and invariants.
3Comprehensive skill library encoding agent-specific guardrails ("always use --dry-run", "always use --fields"). Skills are versioned, discoverable, and follow a standard like OpenClaw.

Interpreting the Total

RangeRatingDescription
0–5Human-onlyBuilt for humans. Agents will struggle with parsing, hallucinate inputs, and lack safety rails.
6–10Agent-tolerantAgents can use it, but they'll waste tokens, make avoidable errors, and require heavy prompt engineering to compensate.
11–15Agent-readySolid agent support. Structured I/O, input validation, and some introspection. A few gaps remain.
16–21Agent-firstPurpose-built for agents. Full schema introspection, comprehensive input hardening, safety rails, and packaged agent knowledge.

Bonus: Multi-Surface Readiness

Not scored, but note whether the CLI exposes multiple agent surfaces from the same binary:

  • MCP (stdio JSON-RPC) — typed tool invocation, no shell escaping
  • Extension / plugin install — agent treats the CLI as a native capability
  • Headless auth — env vars for tokens/credentials, no browser redirect required

google-labs-code의 다른 스킬

typed-service-contracts
google-labs-code
Spec and Handler" 패턴을 사용하여 견고하고 타입 안전한 TypeScript 서비스를 구축하기 위한 아키텍처 표준입니다. CLI, 라이브러리 또는 복잡한…
stitch::code-to-design
google-labs-code
프론트엔드 코드(Vite, React 등)를 정적 HTML 추출, 디자인 시스템 추출, 파일 업로드를 연쇄적으로 수행하여 Stitch Design으로 변환합니다. **항상** 이 방법을 사용하세요…
react-vite-dashboard
google-labs-code
Stitch 디자인을 TanStack Query, DESIGN.md의 접근 가능한 토큰, Web3 준비 패턴(ethers/viem)과 함께 프로덕션 React + Vite 대시보드로 변환합니다.
stitch::extract-design-md
google-labs-code
프론트엔드 소스 코드(React, Vue, Svelte, Angular, 일반 HTML/CSS 또는 모든 웹 프레임워크)로부터 포괄적인 디자인 시스템(DESIGN.md)을 직접 추출합니다.
remotion
google-labs-code
Stitch 앱 디자인에서 Remotion을 사용하여 부드러운 전환과 텍스트 오버레이가 포함된 전문 워크스루 비디오를 제작합니다. Stitch 프로젝트에서 화면을 가져와 확대 효과, 페이드 전환, 상황별 텍스트 오버레이와 함께 Remotion 비디오 구성으로 오케스트레이션합니다. ScreenSlide 및 WalkthroughComposition 컴포넌트를 포함한 모듈식 컴포넌트 아키텍처와 대화형 핫스팟 및 음성 해설 통합과 같은 고급 기능을 지원합니다. 화면 매니페스트를 생성하고 다운로드합니다...
ink
google-labs-code
Ink 터미널 렌더러로, JSON 사양을 대화형 터미널 UI로 변환합니다. @json-render/ink로 작업하거나 터미널 UI를 구축할 때 사용하세요.
tdd-red-green-refactor
google-labs-code
이 스킬은 AI 지원 프로그래밍을 위한 구조적 프레임워크를 구현하여 모든 코드 라인이 검증 가능하고, 타입이 지정되며, 목적에 부합하도록 보장합니다.
automate-github-issues
google-labs-code
병렬 Jules 코딩 에이전트를 사용하여 자동화된 GitHub 이슈 분류 및 해결 설정