fleet-health-report

tarafından nvidia

Canlı nvfleetint verilerinden, düğüm sağlığı, kapasite, aktif uyarı etkisi, son hatalar ve… dahil olmak üzere bağımsız bir filo geneli HTML sağlık anlık görüntüsü oluşturun.

npx skills add https://github.com/nvidia/fleet-intelligence-client --skill fleet-health-report

Fleet Health Report

Generate one offline HTML snapshot from fresh nvfleetint JSON. Read the CLI contract, HTML theme, and workspace guide before collecting data.

Workflow

1. Resolve and verify the profile

Resolve credentials before any report query:

nvfleetint auth list --output json
nvfleetint auth status --profile <profile> --output json

Use the user-named profile or the sole configured profile. If multiple profiles exist and none was requested, ask which one to use and identify the current one as the default suggestion. Require connection equal to ok, then pass the same explicit --profile <profile> to every API-backed command below.

2. Collect overview

Collect the tenant overview before inventory:

nvfleetint overview --profile <profile> --output json

Use it for the entire-fleet headline. In a scoped report, label it fleet-wide context and derive scoped totals from the filtered node list instead.

3. Collect inventory and resolve scope

Accept the entire fleet, compute-zone names, or node-group names. List compute zones first, node groups second, and nodes third. Resolve supplied names to IDs internally; clarify only ambiguous name matches.

Probe each list with the same filters, --view basic where supported, and --page-size 1 before its full pull. Collect the full lists in this order:

nvfleetint computezone list --all --profile <profile> --output json
nvfleetint nodegroup list --all --profile <profile> --output json
nvfleetint node list <scope> --all --profile <profile> --output json

After the first two lists, apply resolved --compute-zone-ids or --nodegroup-ids to the node query.

4. Collect recent errors

Pin one 24-hour error window. Use GNU date -u -d "@$now" and date -u -d "@$((now - 86400))", or BSD/macOS date -u -r "$now" and date -u -r "$now" -v-24H, formatted as RFC3339 UTC.

nvfleetint report error --view list --group-by error \
  --start "$start" --end "$end" --all \
  --profile <profile> --output json

Sum row count; pagination total counts grouped rows, not error occurrences. The error API cannot filter by zone/group, so label it tenant-wide in a scoped report or omit it when strictly scoped evidence is required.

5. Collect filtered active alerts

Discover the server-supported filter values first:

nvfleetint alert options --view active --profile <profile> --output json

From the returned componentTypes options, build a comma-separated list of component IDs excluding exact IDs psirt and agent_liveness. Stop if no component IDs remain.

Request the filtered count of all affected nodes and up to 10 machines ordered by active-alert count:

nvfleetint alert summary <scope> --view active \
  --component-type <component-types> \
  --sort-by alert --order desc --page-size 10 \
  --profile <profile> --output json

Use summary .total for all Nodes with Active Alerts and .totalCritical/.totalWarning for filtered fleet-wide severity totals. The bounded page contains up to 10 machines ordered by active-alert count; state showing N of total when applicable. Do not fetch every affected node merely to count or rank them.

For each returned UUID only, fetch its full filtered drill-down:

nvfleetint alert node <node_uuid> --view active \
  --component-type <component-types> --all \
  --profile <profile> --output json

Run at most four node calls concurrently. Alert collection invokes at most 12 nvfleetint commands: one options command, one summary command, and up to 10 node commands. An --all command may make multiple paginated requests.

6. Derive and write the report

  • Entire fleet: show overview.healthPercentage once as Fleet Health Percentage. Scoped: show 100 * healthy / total once as Healthy Node Percentage and its formula. Do not show both or repeat the percentage in Fleet Summary.
  • Keep node health/severity at source level; present fleet counts and distributions without an aggregate fleet severity badge.
  • Join the returned summary machines to inventory by UUID and preserve the returned order.
  • Use node-alert rows only for those machines' drill-down detail; do not infer unseen component distribution.

Apply the shared HTML theme and use these four report sections:

  1. Fleet Summary: executive summary, fleet totals, and health distribution.
  2. Fleet Distribution: compute zones, node groups, capacity, and operational signals.
  3. Active Alerts: filtered Nodes with Active Alerts, severity aggregates, Machines Needing Immediate Attention (up to 10), and one expandable per-machine drill-down showing component, status, start time, last update, and available evidence text.
  4. Error Distribution: grouped recent errors over the pinned window.

7. Validate and deliver

Apply the CLI contract's completeness checks. The bounded summary page is the only intentional partial result; use its top-level aggregates for fleet-wide claims. For an entire-fleet report, reconcile overview totals with complete inventory totals.

Validate the final HTML:

for id in at-a-glance distribution alerts errors; do
  grep -q "id=\"$id\"" "$out" || exit 1
done
grep -q '</html>' "$out" && ! grep -qE '\[[a-z_]+\]' "$out"

Leave only the final HTML and return its path, scope, collection time, and error window.

nvidia tarafından daha fazla skill

compileiq-debug
nvidia
Bir şeyler yanlış olduğunda kullanın: Search() takılıyor, tüm değerlendirmeler INVALID_SCORE döndürüyor, puanlar iyileşmiyor, her yapılandırma aynı sayıyı döndürüyor, ptxas hataları…
create-github-pr
nvidia
gh CLI kullanarak GitHub pull request'leri oluşturun. Kullanıcı yeni bir PR oluşturmak, kodu incelemeye göndermek veya bir pull request açmak istediğinde kullanın. Tetikleyici anahtar kelimeler -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Diğer açık sorunları tarayarak, belirli bir PR'ın da düzeltebileceği veya yanlışlıkla bozabileceği sorunları bulur. Dosya:satır… ile bitişik düzeltme fırsatlarını ve çelişki risklerini çıktı olarak verir.
fhir-basics
nvidia
Ajanlara FHIR R4 API'lerinin nasıl çalıştığını, hangi kaynakların mevcut olduğunu, arama parametreleriyle nasıl sorgulanacağını ve tüm yanıt formatlarının nasıl doğru şekilde ayrıştırılacağını öğretir…
compileiq-validate-result
nvidia
Bir arama tamamlandıktan SONRA ve herhangi bir hızlandırma talep etmeden veya bir ACF göndermeden ÖNCE kullanın. dump_results CSV dosyasını yükler, en iyi K adayı (tek amaçlı) çıkarır…
changelog-audit
nvidia
Bir sürüm öncesinde Warp CHANGELOG.md dosyasını denetle: kayıp girdileri kurtar, kullanıcı etkisine göre sırala, girdi dilini iyileştir, satır kaydırma yap ve (sürüm dalı modunda) karşılaştırmayı artır…
maintain-dynamic-plugins
nvidia
NeMo Relay dinamik eklenti yükleyicilerini, manifestolarını, Rust yerel SDK'larını, gRPC işçi protokolünü, Python işçi SDK'sını, dokümantasyonu, testleri ve sürüm iş akışı kapsamını korur
dgx-diagnose
nvidia
Yaygın DGX Station GB300 sorunlarını teşhis edin — CUDA çökmeleri, yanlış GPU hedefleme, vLLM/SGLang konteyner hataları, MIG durumu sorunları, NVLink/Fabric Manager hataları,…