k8s-launch-kit-validate

bởi nvidia

Sử dụng kỹ năng này khi người dùng muốn xác minh rằng một triển khai mạng NVIDIA khớp với cấu hình đã tạo ra nó. Kích hoạt cho: 'triển khai của tôi có…

npx skills add https://github.com/nvidia/k8s-launch-kit --skill k8s-launch-kit-validate

l8k: Validate

PREREQUISITE: Read ../k8s-launch-kit-shared/SKILL.md for install paths, global flags, and exit codes.

This workflow uses the default host target; --target host is equivalent. The CLI snapshots validation arguments and runs the standalone Host validation and report service through the target registry; checks and report semantics are unchanged.

Verify that a previously generated and deployed NVIDIA networking deployment is correctly applied and matches the selected Network Operator release.

What it checks

  1. Network Operator Helm release version. Reads the chart's appVersion from any release Secret named sh.helm.release.v1.<release>.v<N> whose release name contains "network-operator", in the operator namespace. Compares with the version expected by networkOperator.selectedRelease in cluster-config.yaml (looked up in l8k's embedded release catalog).
  2. Manifest state. Every YAML manifest under --deployment-files (skipping example workloads and Helm values.yaml) is fetched from the cluster and classified by the per-Kind resource-state registry. NicConfigurationTemplate and NicFirmwareTemplate wait for the operator-populated status.nicDevices list to reflect current node, NIC type, PCI-address, serial-number, and part-number selectors and validate only the named devices. Their propagated template payload and condition observedGeneration must also be current; unrelated discovered NIC configuration state is ignored. A configuration template checks FirmwareUpdateInProgress only for a matched device with spec.firmware, so a stale firmware condition does not block a deployment with no NicFirmwareTemplate. Each manifest is reported READY, IN-PROGRESS, ERROR, or MISSING.
  3. Connectivity matrix. By default, l8k validate applies the generated example DaemonSet, waits for ready pods, and runs source-bound icmp, rping, and ib_write_bw tests. When validation.gpuDirect.enabled is true and ib_write_bw is selected, a distinct DMA-BUF bandwidth family uses the connected GPU index for each source and destination rail. Profile templates declare both validation containers: the release-specific full-runtime DOCA image for the RDMA checks and netshoot for ICMP and route checks. Validate applies the generated DaemonSet without runtime container injection. The default mode is strict.

Validation also reports deploy-preflight drift without remediating it. SriovNetworkPoolConfig, SriovNetworkNodePolicy, and OVSNetwork objects labeled with spectrumx.nvidia.com/owner-name are excluded from stray results because the Spectrum-X operator owns them as children of SpectrumXRailPoolConfig.

Exit code is non-zero (4) on any missing manifest, version mismatch, or gating connectivity failure. Version checks soft-skip when prerequisites are absent — no cluster-config.yaml, no Helm release Secret, etc.

Usage

l8k validate [--user-config <PATH>] [--deployment-files <DIR>] [--kubeconfig <PATH>]

Flags

FlagDefaultDescription
--kubeconfig$KUBECONFIGPath to kubeconfig with read access to the cluster
--user-config./cluster-config.yamlCluster config YAML; used for networkOperator.selectedRelease and the operator namespace
--deployment-files./deploymentDirectory containing the manifests to verify
--validation-modevalidation.mode (strict)Connectivity mode: quick, full, or strict
--validation-checksvalidation.checks (icmp,rping,ib_write_bw)Comma-separated connectivity checks; "" disables all
--connectivity-timeoutautomatic (0)Total connectivity budget is calculated from the generated matrix plan; set a positive duration for an explicit hard setup and execution deadline
--rdma-rping-iterationsvalidation.rdma.rpingIterationsrping client iteration count
--rdma-ib-write-sizevalidation.rdma.ibWriteSizeib_write_bw message size
--rdma-ib-write-min-bandwidth-gbpsvalidation.rdma.ibWriteMinBandwidthGbpsMinimum peak Gbps; 0 disables bandwidth gating
--log-leveldisableddebug for structured progress and timing; trace also includes bounded command output

Connectivity Modes

  • quick: all same-rail node pairs plus one non-gating cross-rail canary per source-rail/destination-rail mapping.
  • full: every source rail × every destination rail × every ordered pod pair; cross-rail results are reported but do not gate pass/fail.
  • strict: full matrix. Cross-rail gates by profile.routing: source-based must succeed, destination-based must stay isolated.

All checks are source-bound. ICMP always uses ping -I <src-iface> for same-rail and cross-rail probes in every validation mode; it never binds only the source IP. rping uses -I <src-ip>, and ib_write_bw uses --bind_source_ip <src-ip>. GPUDirect adds --use_cuda=<endpoint-index> --use_cuda_dmabuf independently on the client and server. Treat missing or ambiguous connectedGPU topology as a failure; never substitute GPU 0. Pull Secrets come from networkOperator.imagePullSecrets and must exist in every validation namespace.

With the default --connectivity-timeout=0, validate logs one total automatic budget after planning the matrix. The calculation reflects the enabled test families, per-command limits, ordered pod-pair batches, bounded workload setup, cleanup, and a safety margin. A positive flag value replaces that calculation with a user-supplied hard deadline.

Examples

# Defaults: ./cluster-config.yaml + ./deployment, $KUBECONFIG
l8k validate

# Explicit paths
l8k validate --user-config ./cluster-config.yaml \
  --deployment-files ./deployment \
  --kubeconfig ~/.kube/config

# Agent mode (single JSON object on stdout, logs on stderr)
l8k validate --output json 2>/dev/null | jq '.summary'

# Diagnose stage or batch progress without raw command output
l8k validate --log-level debug

# Capture bounded route, ICMP, RDMA client, and RDMA server evidence
l8k validate --log-level trace --keep

Debug logs show the endpoint inventory, plan, source-route cache statistics, static checks, stages, RDMA batches, cleanup, report writes, elapsed time, and remaining timeout. Trace adds bounded commands and per-test stdout/stderr. Failed RDMA server logs are collected before the temporary files and test workload are removed. Add --keep only when follow-up pod inspection is needed.

Output

Text mode prints a short report:

Network Operator release
  selectedRelease: 26.4
  expected version: v26.4.0-beta.6
  deployed: network-operator (chart=26.4.0-beta.6 app=v26.4.0-beta.6 rev=3 status=deployed)
  result: MATCH

Manifests
  [READY      ] NicClusterPolicy/nic-cluster-policy in (cluster-scoped)
  [IN-PROGRESS] NicConfigurationTemplate/spectrum-x-config in network-operator — waiting for nic-configuration-operator to populate status.nicDevices with matched devices
  [MISSING    ] SriovNetwork/sriov-network-rail-0 in default — not found in cluster
  ...

Summary: 1/3 ready, 1 in-progress, 0 error, 1 missing; version: match; topology mismatches: 0 group(s)

JSON mode (--output json) emits one object with versionCheck, manifests, and summary fields.

When this skill activates

Trigger phrases include: "validate my deployment", "is my cluster correct", "are all the manifests applied", "does the chart version match", "did the deploy succeed", or any discrepancy claim about expected vs deployed state.

See Also

Thêm skills từ nvidia

compileiq-debug
nvidia
Sử dụng khi có điều gì đó không ổn: Search() bị treo, tất cả các đánh giá đều trả về INVALID_SCORE, điểm số không cải thiện, mọi cấu hình đều trả về cùng một số, lỗi ptxas…
create-github-pr
nvidia
Tạo pull request GitHub bằng cách sử dụng gh CLI. Sử dụng khi người dùng muốn tạo PR mới, gửi mã để xem xét, hoặc mở pull request. Từ khóa kích hoạt -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Quét các vấn đề đang mở khác để tìm những vấn đề mà một PR nhất định có thể sửa hoặc vô tình làm hỏng. Đưa ra các cơ hội sửa lỗi liền kề và rủi ro mâu thuẫn với file:dòng…
fhir-basics
nvidia
Dạy các tác nhân cách hoạt động của API FHIR R4, những tài nguyên có sẵn, cách truy vấn chúng với tham số tìm kiếm, và cách phân tích chính xác tất cả các định dạng phản hồi…
compileiq-validate-result
nvidia
Sử dụng SAU KHI tìm kiếm hoàn tất và TRƯỚC KHI yêu cầu tăng tốc hoặc gửi ACF. Tải tệp CSV dump_results, trích xuất các ứng viên top-K (đơn mục tiêu)…
changelog-audit
nvidia
Kiểm tra Warp CHANGELOG.md trước khi phát hành: khôi phục các mục bị mất, sắp xếp theo tác động người dùng, tinh chỉnh ngôn ngữ mục, xuống dòng và (chế độ nhánh phát hành) so sánh bump…
maintain-dynamic-plugins
nvidia
Duy trì các bộ nạp plugin động NeMo Relay, tệp kê khai, SDK gốc Rust, giao thức worker gRPC, SDK worker Python, tài liệu, kiểm thử và phạm vi quy trình phát hành
dgx-diagnose
nvidia
Chẩn đoán các sự cố thường gặp của DGX Station GB300 — lỗi CUDA, nhắm sai GPU, lỗi container vLLM/SGLang, vấn đề trạng thái MIG, lỗi NVLink/Fabric Manager,…