aicr-auditing-docs

โดย nvidia

ใช้เมื่อตรวจสอบเอกสาร Markdown ของ AICR ว่ามีการซ้ำซ้อน คลาดเคลื่อน เนื้อหาเกินจำเป็น และช่องว่าง — เพื่อรักษาเอกสารให้มีคุณค่าสูงในขณะที่โปรเจกต์พัฒนาไป เริ่มทำงานเมื่อมีคำว่า "audit…

npx skills add https://github.com/nvidia/aicr --skill aicr-auditing-docs

Auditing AICR Documentation

Overview

AICR's docs are mature and already well-structured: a persona split under docs/ (user / integrator / contributor), a docs hub with glossary (docs/README.md), ADRs in docs/design/, and a strong root README.md. The recurring risk is not missing structure — it is duplication and drift as features land. This skill is a repeatable audit, not a rewrite: default to a findings report; only edit when explicitly asked.

When to Use

  • Periodic doc health check, or before a release.
  • After a large feature merges (new CLI flag, API endpoint, component, recipe field).
  • When the same explanation appears in multiple files and you suspect drift.
  • Not for: writing a single new doc (just write it, following the style rules below), or generated content (docs/conformance/, docs/user/container-images.md).

The Map (what to audit, and how it's owned)

AreaCanonical ownerNotes
Project pitch, features, supported envsroot README.mdQuick Start may duplicate docs/user/installation.md — acceptable for README only.
User how-to / referencedocs/user/cli-reference.md owns flags; task narrative belongs in task docs (validation.md, agent-deployment.md).
Integration / embeddingdocs/integrator/Resolver internals belong in contributor/, not here.
Project internalsdocs/contributor/Architecture overview = contributor/index.md.
Demos / runbooksdemos/GitHub-only (not in fern/docs.yml nav). Terse names hurt discovery.
GovernanceCONTRIBUTING.md, DEVELOPMENT.md, RELEASING.md, SECURITY.mdEach owns one concern; cross-link instead of repeating.
Agent rules.claude/CLAUDE.md (canonical) → AGENTS.md (CI-synced mirror — never flag).github/copilot-instructions.md should be a pointer, not a copy.

Sources of Truth (drift hotspots — check these first)

Drift between examples and these authoritative sources is the highest-value class of finding:

  • Component/chart/image versionsdocs/user/container-images.md (the BOM, regenerated by make bom-docs). Inline version examples elsewhere (cli-reference.md, api-reference.md, data-flow.md) should be marked illustrative and point here — never hand-pinned to a stale tag.
  • Tool versions (golangci-lint, Go, Ko) → .settings.yaml. Never hardcode in prose or sample workflow YAML.
  • Criteria/enum values (service, accelerator, os, intent, platform, error codes) → the Go type (e.g. pkg/recipe/criteria.go). Enums are enumerated in many files; see the enum-audit checklist in CLAUDE.md → Documentation updates.
  • API shapeapi/aicr/v1/server.yaml (OpenAPI). The component lists and response samples in api-reference.md drift from the registry — diff them.

The Five Audit Dimensions

Run every file through these. Group findings by dimension, prioritized.

  1. Duplication — same steps/tables/concepts in >1 file. Pick the canonical owner (table above), trim the rest to a cross-link. Known repeat offenders: agent/snapshot deployment, recipe-evidence walkthrough, constraint paths/operators table, make-target block, DCO/signing rules, the six demos/cuj*-{eks,gke}.md files (~80% shared).
  2. Overlap / misplacement — content in the wrong persona tree (e.g., resolver internals in integrator/, agent deployment in integrator/ when it's a user concern), or a topic split awkwardly across files.
  3. Bloat — verbose sections that lose no value when cut: embedded SBOM/JSON dumps, repeated export TAG= blocks, stub code that contradicts the architecture ("AICR is not a controller"), generic K8s boilerplate that isn't AICR-specific, walls of text needing structure.
  4. Gaps (Diátaxis) — is each of tutorial / how-to / reference / explanation present for the persona? Known holes: no end-to-end tutorial (install→recipe→bundle→deploy→validate), no aicr bundle how-to.
  5. Staleness / style violations — TODOs, "(Future)" content shown as usable, contradictory version numbers, leftover template placeholders (__AICR (AICR)__), and violations of the repo's own doc-style rules below.

Repo Doc-Style Rules (enforce during audit)

These are defined in CLAUDE.mdDocumentation Style; flag violations:

  • Auto-anchors, no manual TOCs — GitHub/Fern generate anchors. Manual ## Table of Contents blocks (present in CONTRIBUTING.md, DEVELOPMENT.md) are violations.
  • Promote **Bold Label:** to a heading sparingly — only a named topic with ≥ ~8 lines beneath it.
  • Anchor hygiene — when renaming/removing a heading, grep <file>.md#<old-slug> repo-wide and fix inbound links. Broken anchors fail CI via lychee on any docs/** PR (.github/workflows/fern-docs-ci.yaml) — but not make qualify.

How to Run It

  1. Parallelize by area to protect context: dispatch one research agent per tree — (a) docs/ persona trees, (b) root + agent docs, (c) demos/. Give each the five dimensions and the sources-of-truth list; ask for a concise (<600 word) report with file paths and concrete recommendations.
  2. Synthesize into one prioritized report. Lead with version/enum drift (highest value), then duplication consolidations, then bloat/gaps.
  3. Report, don't rewrite by default. If asked to fix: one focused PR per theme (e.g., "de-dup agent deployment", "trim SECURITY.md"), keep diffs reviewable, and after any change run make qualify (and note lychee is separate). Touching docs/** requires the lychee anchor check.

Common Mistakes

  • Flagging the AGENTS.mdCLAUDE.md mirror as duplication — it's intentional and CI-enforced.
  • Recommending a TOC "for navigation" — violates the auto-anchor rule.
  • Hand-fixing a version in an example — fix the pattern (mark illustrative, link to container-images.md) so it can't re-drift.
  • Editing generated files (container-images.md, docs/conformance/) directly instead of their generators.
  • Auditing one file in isolation — duplication only shows up across files; read the tree.

Skills เพิ่มเติมจาก nvidia

compileiq-debug
nvidia
ใช้เมื่อมีบางอย่างผิดปกติ: Search() ค้าง, การประเมินทั้งหมดคืนค่า INVALID_SCORE, คะแนนไม่ดีขึ้น, ทุกคอนฟิกคืนค่าเลขเดียวกัน, ข้อผิดพลาด ptxas…
create-github-pr
nvidia
สร้างคำขอดึงข้อมูล GitHub โดยใช้ gh CLI ใช้เมื่อผู้ใช้ต้องการสร้าง PR ใหม่ ส่งโค้ดเพื่อตรวจสอบ หรือเปิดคำขอดึงข้อมูล คำหลักที่ใช้เรียก -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
สแกน issue อื่นๆ ที่เปิดอยู่เพื่อค้นหาว่า PR ที่กำหนดอาจแก้ไขหรือทำให้เสียโดยไม่ได้ตั้งใจ แสดงผลโอกาสในการแก้ไขที่เกี่ยวข้องและความเสี่ยงที่ขัดแย้งกันพร้อมไฟล์:บรรทัด…
fhir-basics
nvidia
สอนให้เอเจนต์เข้าใจการทำงานของ FHIR R4 API ทรัพยากรที่มีให้ วิธีค้นหาด้วยพารามิเตอร์ค้นหา และวิธีแยกวิเคราะห์รูปแบบการตอบกลับทั้งหมดอย่างถูกต้อง…
compileiq-validate-result
nvidia
ใช้หลังจากที่การค้นหาเสร็จสิ้น และก่อนที่จะอ้างสิทธิ์การเร่งความเร็วหรือจัดส่ง ACF โหลดไฟล์ CSV dump_results แยกผู้สมัคร K อันดับแรก (วัตถุประสงค์เดียว)…
changelog-audit
nvidia
ตรวจสอบ Warp CHANGELOG.md ก่อนปล่อย: กู้คืนรายการที่สูญหาย จัดเรียงตามผลกระทบต่อผู้ใช้ ปรับปรุงภาษาในรายการ จัดบรรทัด และ (ในโหมดสาขาปล่อย) เปรียบเทียบการเพิ่มเวอร์ชัน…
maintain-dynamic-plugins
nvidia
ดูแล NeMo Relay dynamic plugin loaders, manifests, Rust native SDKs, gRPC worker protocol, Python worker SDK, เอกสาร, การทดสอบ และความครอบคลุมของเวิร์กโฟลว์การเผยแพร่
dgx-diagnose
nvidia
วินิจฉัยปัญหาทั่วไปของ DGX Station GB300 — CUDA ล่ม, การกำหนดเป้าหมาย GPU ผิด, บั๊กคอนเทนเนอร์ vLLM/SGLang, ปัญหาสถานะ MIG, ข้อผิดพลาด NVLink/Fabric Manager,…