nvcf-explore-stack

por nvidia

Navega y explica el stack autoalojado de NVCF dentro del monorepo. Mapea los releases de helmfile con sus charts, subárboles de fuentes de imágenes, hooks de helm, namespaces, y…

npx skills add https://github.com/nvidia/nvcf --skill nvcf-explore-stack

NVCF Explore Stack

Help a user or developer navigate the NVCF self-hosted stack inside the monorepo. Identify what a release is, what installs it, what it depends on, and which subtree owns the chart and the runtime image.

Instructions

This skill is for both NVCF users and NVCF developers. Users use it to understand stack dependencies, deployment order, and CI/CD migration questions. Developers use it to find the owning chart, source subtree, hook, or helmfile stage before making changes.

Use this skill long enough to answer the question, then hand off to the right execution skill. Always cite the source file path so the user can verify.

Required inputs

Read these from the monorepo root.

Authoritative (always read first when answering):

  • deploy/stacks/self-managed/helmfile.d/00-observability-infrastructure.yaml.gotmpl
  • deploy/stacks/self-managed/helmfile.d/01-dependencies.yaml.gotmpl
  • deploy/stacks/self-managed/helmfile.d/02-core.yaml.gotmpl
  • deploy/stacks/self-managed/helmfile.d/03-observability.yaml.gotmpl
  • deploy/stacks/observability/helmfile.d/01-observability.yaml.gotmpl
  • deploy/stacks/nvcf-compute-plane/helmfile.d/01-dependencies.yaml.gotmpl
  • deploy/stacks/nvcf-compute-plane/helmfile.d/02-nvca.yaml.gotmpl

Chart-level (when the chart is checked into the monorepo):

  • deploy/helm/<chart>/Chart.yaml
  • deploy/helm/<chart>/values.yaml

Common questions

What deploys X : Look up release X in the helmfile stage files. Return chart name, version, namespace, and which gotmpl file declares it. If the chart is checked in, also point at deploy/helm/<chart>/.

What does X depend on : Return the needs: chain for that release plus the stage gate it sits behind (control-plane stages 0 -> 1 -> 2 -> 3, then compute-plane). Include any profile, condition:, or component mode that gates whether X deploys at all.

What hooks run for X : Read the checked-in chart under deploy/helm/<chart>/ when available. Search its templates/ directory for Helm hook annotations, weights, hook events (pre-install / post-install), images used, and purpose. Cite the chart-relative template file path. If the chart is not checked in, cite its Helmfile chart reference and state that local templates are unavailable.

Walk me through the full deployment order : Summarize control-plane stages 0 through 3 from the self-managed gotmpl files. Stage 0 delegates to the shared observability Helmfile when observability.profile is enabled. Stage 3 installs State Metrics and the function autoscaler for control and all. Then summarize the compute-plane stage from deploy/stacks/nvcf-compute-plane/helmfile.d/01-dependencies.yaml.gotmpl and deploy/stacks/nvcf-compute-plane/helmfile.d/02-nvca.yaml.gotmpl. Call out which releases run in parallel inside a stage and which are serialized by needs:.

Which subtree do I edit to change X : Point at the Helmfile path for orchestration, deploy/helm/<chart>/ for checked-in chart wiring, and the image source under src/, infra/, or migrations/ for runtime behavior. If a referenced chart is not checked in, report its Helmfile chart reference and do not claim a local source path.

What namespaces does the stack use : Return the list from the helmfile (namespace: per release).

Subtree mapping

The stack lives in three layers across the monorepo:

ConcernLives at
Helmfile orchestration (stage ordering, env wiring, secrets flow)deploy/stacks/self-managed/
Chart manifests, helm hooks, valuesdeploy/helm/<chart>/ when checked in; otherwise use the Helmfile chart reference
Runtime application code, migrationssrc/, infra/, migrations/

When a question crosses layers, answer by layer and tell the user the order to edit (chart wiring first if the deploy contract changes, image source if behavior changes).

Tone

Assume the user is onboarding to the stack. Be concise. Always include the chart name, version, and the gotmpl path when referencing a release. Prefer one short paragraph plus a code-block citation over prose.

Skill handoff candidates

After exploring, suggest the next skill when applicable:

  • nvcf-self-managed-installation for installing, upgrading, or tearing down the stack
  • docs/dev/local-development.md for k3d / local cluster work
  • nvcf-self-managed-cli for nvcf-cli usage against an installed stack
  • docs/AGENTS.md and fern/versions/main.yml for routing the user to a published docs page
  • tools/ci/check-doc-version-sync for keeping the documentation manifest in sync with the docs version catalog

Más skills de nvidia

compileiq-debug
nvidia
Úsalo cuando algo esté mal: Search() se cuelga, todas las evaluaciones devuelven INVALID_SCORE, las puntuaciones no mejoran, cada configuración devuelve el mismo número, errores de ptxas…
create-github-pr
nvidia
Crear solicitudes de extracción de GitHub usando la CLI gh. Usar cuando el usuario quiera crear un nuevo PR, enviar código para revisión o abrir una solicitud de extracción. Palabras clave de activación -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Escanea otros issues abiertos para encontrar aquellos que un PR dado también podría corregir o romper accidentalmente. Genera oportunidades de corrección adyacente y riesgos de contradicción con archivo:línea…
fhir-basics
nvidia
Enseña a los agentes cómo funcionan las APIs de FHIR R4, qué recursos están disponibles, cómo consultarlos con parámetros de búsqueda y cómo analizar correctamente todos los formatos de respuesta…
compileiq-validate-result
nvidia
Usar DESPUÉS de que una Búsqueda haya finalizado y ANTES de reclamar cualquier aceleración o enviar un ACF. Carga el CSV de dump_results, extrae los mejores K candidatos (de un solo objetivo)…
changelog-audit
nvidia
Auditar el CHANGELOG.md de Warp antes de un lanzamiento: recuperar entradas perdidas, ordenar por impacto en el usuario, refinar el lenguaje de las entradas, ajustar saltos de línea y (en modo rama de lanzamiento) incrementar comparación…
maintain-dynamic-plugins
nvidia
Mantener los cargadores de plugins dinámicos de NeMo Relay, manifiestos, SDKs nativos de Rust, protocolo de trabajador gRPC, SDK de trabajador Python, documentación, pruebas y cobertura del flujo de trabajo de lanzamiento
dgx-diagnose
nvidia
Diagnostica problemas comunes de la DGX Station GB300: fallos de CUDA, direccionamiento incorrecto de GPU, errores de contenedores vLLM/SGLang, problemas de estado MIG, errores de NVLink/Fabric Manager,…