k8s-launch-kit-troubleshoot
Gunakan keterampilan ini ketika pengguna mengalami masalah dengan NVIDIA Network Operator di Kubernetes, atau ingin menganalisis dump diagnostik sosreport. Aktifkan untuk: OFED…
npx skills add https://github.com/nvidia/k8s-launch-kit --skill k8s-launch-kit-troubleshootl8k: Troubleshooting
PREREQUISITE: Read
../k8s-launch-kit-shared/SKILL.mdfor install paths, global flags, and exit codes.
Debug NVIDIA Network Operator issues on Kubernetes, with or without a sosreport.
l8k Troubleshooting Commands
# Show validation endpoints, planning, route-cache statistics, stage/batch
# progress, timings, and failed RDMA evidence.
l8k validate --kubeconfig <PATH> --deployment-files <DIR> --log-level debug
# Add bounded route command output, RDMA client stdout/stderr, and server logs.
l8k validate --kubeconfig <PATH> --deployment-files <DIR> --log-level trace --keep
# Collect a diagnostic dump from the cluster
l8k sosreport [--kubeconfig <PATH>] --output-dir ./sosreport
Start with debug. Escalate to trace when a route, ICMP, rping, or
ib_write_bw failure needs command output. Trace fields are bounded; failed RDMA
server logs are collected before cleanup. Add --keep only when the workload
must remain available for follow-up kubectl exec inspection.
l8k sosreport gathers cluster state, CRDs, operator logs, and per-node NIC info into a structured directory for offline analysis. Use the diagnostic commands and triage workflow below to interpret the dump.
Diagnostic Commands
# NicClusterPolicy status (check state: ready vs notReady)
kubectl get nicclusterpolicy -o yaml
# Network operator pods
kubectl get pods -n <operator-ns> -o wide
# SR-IOV node states (VF allocation)
kubectl get sriovnetworknodestates -A -o yaml
# OFED driver pod logs
kubectl logs -n <operator-ns> -l app=mofed-<os> --tail=100
# NIC configuration daemon logs
kubectl logs -n <operator-ns> -l app=nic-configuration-daemon --tail=100
# Check for pods stuck on network resources
kubectl get pods -A -o wide | grep -E 'ContainerCreating|Init'
Common Failure Patterns
| Symptom | Likely Cause | Fix |
|---|---|---|
NicClusterPolicy state: notReady | OFED driver pods failing | Check mofed pod logs, verify kernel/driver compatibility |
Pods stuck in ContainerCreating | VFs not allocated or SR-IOV policy not applied | Check sriovnetworknodestates, verify device plugin pods |
CrashLoopBackOff on mofed pods | Kernel module conflict | Check thirdPartyRDMAModules, enable unloadThirdPartyRDMAModules |
| No VFs on node | SriovNetworkNodePolicy not matching | Verify nodeSelector labels match worker nodes |
| RDMA not working | Missing RDMA device plugin or wrong resource name | Check rdma-shared-dp pods, verify resource annotations |
| Phase 0 Helm chart download returns HTTP 401 or an image-pull-Secret error | The configured Secret is missing from the operator namespace, unreadable by the kubeconfig, or has no compatible Docker auth entry | Verify the Secret in networkOperator.namespace; for NGC, its .dockerconfigjson must contain nvcr.io credentials |
l8k discover daemon pods stuck (ImagePullBackOff / Pending) | Bad image tag, missing pull secret, or no feature.node.kubernetes.io/pci-15b3.present=true nodes | Re-run with --keep-namespace then kubectl describe pod -n nvidia-k8s-launch-kit. Fix networkOperator.componentVersion / pass --image-pull-secrets / verify NFD is running. |
l8k validate / deploy can't find Network Operator pods | Operator namespace mismatch | Verify --network-operator-namespace matches actual namespace (does NOT apply to l8k discover — it ignores the flag and uses its own nvidia-k8s-launch-kit namespace) |
| IPPool not allocating | NV-IPAM subnet exhausted or misconfigured | Check ippools CR status, verify CIDR ranges |
--for requires --node-selector | --for was passed without --node-selector | Add --node-selector key=val,…. The synthesized clusterConfig has no live worker-node list; the selector identifies target nodes at apply time. |
--for and --discover-cluster-config are mutually exclusive | Both flags passed simultaneously | Pick one: --for skips discovery, --discover-cluster-config runs it. |
unknown preset "X"; available: … | --for X doesn't match any directory under presets/ | Run l8k preset list and re-run with one of those names. |
preset has no capabilities block | Preset YAML used by --for is missing capabilities.nodes.{sriov,rdma,ib} | Add the block to the preset's topology.yaml. Discovery-time overlay does not require it; only --for does. |
unknown field "productType" in YAML | Hand-authored config still uses the old key name | Rename productType: to gpuType: (the field was renamed). |
For detailed triage workflow, read references/troubleshooting-guide.md.
sosreport Analysis
If the user has a pre-collected sosreport directory (from l8k sosreport or manual collection):
sosreport/
├── metadata/ # Cluster info, node list
├── crds/ # NicClusterPolicy, SriovNetworkNodePolicy, IPPool, etc.
├── operator/ # Network operator pod logs
├── nodes/ # Per-node device info
└── network/ # Interface config, routing tables
Triage Checklist
- Read
metadata/diagnostic-summary.yamlfor overview - Check pod health in
operator/pods.yaml - Inspect CRDs in
crds/for status fields - Read operator logs in
operator/logs/for errors - Check per-node NIC state in
nodes/<node>/
See Also
- k8s-launch-kit-shared — Exit codes and error structure
- k8s-launch-kit-discover — Re-discover to verify hardware state
references/troubleshooting-guide.md— Detailed triage workflow