aicr-auditing-docs

por nvidia

Use ao revisar a documentação Markdown da AICR para verificar duplicação, desvios, inchaço e lacunas — para manter a documentação de alto valor conforme o projeto evolui. Aciona em "auditoria…

npx skills add https://github.com/nvidia/aicr --skill aicr-auditing-docs

Auditing AICR Documentation

Overview

AICR's docs are mature and already well-structured: a persona split under docs/ (user / integrator / contributor), a docs hub with glossary (docs/README.md), ADRs in docs/design/, and a strong root README.md. The recurring risk is not missing structure — it is duplication and drift as features land. This skill is a repeatable audit, not a rewrite: default to a findings report; only edit when explicitly asked.

When to Use

  • Periodic doc health check, or before a release.
  • After a large feature merges (new CLI flag, API endpoint, component, recipe field).
  • When the same explanation appears in multiple files and you suspect drift.
  • Not for: writing a single new doc (just write it, following the style rules below), or generated content (docs/conformance/, docs/user/container-images.md).

The Map (what to audit, and how it's owned)

AreaCanonical ownerNotes
Project pitch, features, supported envsroot README.mdQuick Start may duplicate docs/user/installation.md — acceptable for README only.
User how-to / referencedocs/user/cli-reference.md owns flags; task narrative belongs in task docs (validation.md, agent-deployment.md).
Integration / embeddingdocs/integrator/Resolver internals belong in contributor/, not here.
Project internalsdocs/contributor/Architecture overview = contributor/index.md.
Demos / runbooksdemos/GitHub-only (not in fern/docs.yml nav). Terse names hurt discovery.
GovernanceCONTRIBUTING.md, DEVELOPMENT.md, RELEASING.md, SECURITY.mdEach owns one concern; cross-link instead of repeating.
Agent rules.claude/CLAUDE.md (canonical) → AGENTS.md (CI-synced mirror — never flag).github/copilot-instructions.md should be a pointer, not a copy.

Sources of Truth (drift hotspots — check these first)

Drift between examples and these authoritative sources is the highest-value class of finding:

  • Component/chart/image versionsdocs/user/container-images.md (the BOM, regenerated by make bom-docs). Inline version examples elsewhere (cli-reference.md, api-reference.md, data-flow.md) should be marked illustrative and point here — never hand-pinned to a stale tag.
  • Tool versions (golangci-lint, Go, Ko) → .settings.yaml. Never hardcode in prose or sample workflow YAML.
  • Criteria/enum values (service, accelerator, os, intent, platform, error codes) → the Go type (e.g. pkg/recipe/criteria.go). Enums are enumerated in many files; see the enum-audit checklist in CLAUDE.md → Documentation updates.
  • API shapeapi/aicr/v1/server.yaml (OpenAPI). The component lists and response samples in api-reference.md drift from the registry — diff them.

The Five Audit Dimensions

Run every file through these. Group findings by dimension, prioritized.

  1. Duplication — same steps/tables/concepts in >1 file. Pick the canonical owner (table above), trim the rest to a cross-link. Known repeat offenders: agent/snapshot deployment, recipe-evidence walkthrough, constraint paths/operators table, make-target block, DCO/signing rules, the six demos/cuj*-{eks,gke}.md files (~80% shared).
  2. Overlap / misplacement — content in the wrong persona tree (e.g., resolver internals in integrator/, agent deployment in integrator/ when it's a user concern), or a topic split awkwardly across files.
  3. Bloat — verbose sections that lose no value when cut: embedded SBOM/JSON dumps, repeated export TAG= blocks, stub code that contradicts the architecture ("AICR is not a controller"), generic K8s boilerplate that isn't AICR-specific, walls of text needing structure.
  4. Gaps (Diátaxis) — is each of tutorial / how-to / reference / explanation present for the persona? Known holes: no end-to-end tutorial (install→recipe→bundle→deploy→validate), no aicr bundle how-to.
  5. Staleness / style violations — TODOs, "(Future)" content shown as usable, contradictory version numbers, leftover template placeholders (__AICR (AICR)__), and violations of the repo's own doc-style rules below.

Repo Doc-Style Rules (enforce during audit)

These are defined in CLAUDE.mdDocumentation Style; flag violations:

  • Auto-anchors, no manual TOCs — GitHub/Fern generate anchors. Manual ## Table of Contents blocks (present in CONTRIBUTING.md, DEVELOPMENT.md) are violations.
  • Promote **Bold Label:** to a heading sparingly — only a named topic with ≥ ~8 lines beneath it.
  • Anchor hygiene — when renaming/removing a heading, grep <file>.md#<old-slug> repo-wide and fix inbound links. Broken anchors fail CI via lychee on any docs/** PR (.github/workflows/fern-docs-ci.yaml) — but not make qualify.

How to Run It

  1. Parallelize by area to protect context: dispatch one research agent per tree — (a) docs/ persona trees, (b) root + agent docs, (c) demos/. Give each the five dimensions and the sources-of-truth list; ask for a concise (<600 word) report with file paths and concrete recommendations.
  2. Synthesize into one prioritized report. Lead with version/enum drift (highest value), then duplication consolidations, then bloat/gaps.
  3. Report, don't rewrite by default. If asked to fix: one focused PR per theme (e.g., "de-dup agent deployment", "trim SECURITY.md"), keep diffs reviewable, and after any change run make qualify (and note lychee is separate). Touching docs/** requires the lychee anchor check.

Common Mistakes

  • Flagging the AGENTS.mdCLAUDE.md mirror as duplication — it's intentional and CI-enforced.
  • Recommending a TOC "for navigation" — violates the auto-anchor rule.
  • Hand-fixing a version in an example — fix the pattern (mark illustrative, link to container-images.md) so it can't re-drift.
  • Editing generated files (container-images.md, docs/conformance/) directly instead of their generators.
  • Auditing one file in isolation — duplication only shows up across files; read the tree.

Mais skills de nvidia

compileiq-debug
nvidia
Use quando algo está errado: Search() trava, todas as avaliações retornam INVALID_SCORE, as pontuações não estão melhorando, toda configuração retorna o mesmo número, erros de ptxas…
create-github-pr
nvidia
Crie pull requests do GitHub usando a CLI gh. Use quando o usuário quiser criar um novo PR, enviar código para revisão ou abrir um pull request. Palavras-chave de acionamento -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Escaneia outras issues abertas para encontrar aquelas que um determinado PR pode também corrigir ou quebrar acidentalmente. Gera oportunidades de correção adjacentes e riscos de contradição com arquivo:linha…
fhir-basics
nvidia
Ensina aos agentes como funcionam as APIs FHIR R4, quais recursos estão disponíveis, como consultá-los com parâmetros de busca e como analisar corretamente todos os formatos de resposta…
compileiq-validate-result
nvidia
Use APÓS a conclusão de uma Pesquisa e ANTES de reivindicar qualquer aceleração ou enviar um ACF. Carrega o CSV dump_results, extrai os K melhores candidatos (objetivo único)…
changelog-audit
nvidia
Auditar o CHANGELOG.md do Warp antes de um lançamento: recuperar entradas perdidas, ordenar por impacto ao usuário, refinar a linguagem das entradas, ajustar quebras de linha e (no modo de branch de lançamento) incrementar comparação…
maintain-dynamic-plugins
nvidia
Manter carregadores de plugins dinâmicos do NeMo Relay, manifestos, SDKs nativos em Rust, protocolo de worker gRPC, SDK de worker Python, documentação, testes e cobertura do fluxo de lançamento
dgx-diagnose
nvidia
Diagnostique problemas comuns do DGX Station GB300 — falhas de CUDA, direcionamento incorreto de GPU, bugs de contêiner vLLM/SGLang, problemas de estado MIG, erros de NVLink/Fabric Manager,…