k8s-launch-kit-troubleshoot

작성자: nvidia

사용자가 Kubernetes에서 NVIDIA Network Operator에 문제가 있거나 sosreport 진단 덤프를 분석하려 할 때 이 스킬을 사용하세요. 활성화 대상: OFED…

npx skills add https://github.com/nvidia/k8s-launch-kit --skill k8s-launch-kit-troubleshoot

l8k: Troubleshooting

PREREQUISITE: Read ../k8s-launch-kit-shared/SKILL.md for install paths, global flags, and exit codes.

Debug NVIDIA Network Operator issues on Kubernetes, with or without a sosreport.

l8k Troubleshooting Commands

# Show validation endpoints, planning, route-cache statistics, stage/batch
# progress, timings, and failed RDMA evidence.
l8k validate --kubeconfig <PATH> --deployment-files <DIR> --log-level debug

# Add bounded route command output, RDMA client stdout/stderr, and server logs.
l8k validate --kubeconfig <PATH> --deployment-files <DIR> --log-level trace --keep

# Collect a diagnostic dump from the cluster
l8k sosreport [--kubeconfig <PATH>] --output-dir ./sosreport

Start with debug. Escalate to trace when a route, ICMP, rping, or ib_write_bw failure needs command output. Trace fields are bounded; failed RDMA server logs are collected before cleanup. Add --keep only when the workload must remain available for follow-up kubectl exec inspection.

l8k sosreport gathers cluster state, CRDs, operator logs, and per-node NIC info into a structured directory for offline analysis. Use the diagnostic commands and triage workflow below to interpret the dump.

Diagnostic Commands

# NicClusterPolicy status (check state: ready vs notReady)
kubectl get nicclusterpolicy -o yaml

# Network operator pods
kubectl get pods -n <operator-ns> -o wide

# SR-IOV node states (VF allocation)
kubectl get sriovnetworknodestates -A -o yaml

# OFED driver pod logs
kubectl logs -n <operator-ns> -l app=mofed-<os> --tail=100

# NIC configuration daemon logs
kubectl logs -n <operator-ns> -l app=nic-configuration-daemon --tail=100

# Check for pods stuck on network resources
kubectl get pods -A -o wide | grep -E 'ContainerCreating|Init'

Common Failure Patterns

SymptomLikely CauseFix
NicClusterPolicy state: notReadyOFED driver pods failingCheck mofed pod logs, verify kernel/driver compatibility
Pods stuck in ContainerCreatingVFs not allocated or SR-IOV policy not appliedCheck sriovnetworknodestates, verify device plugin pods
CrashLoopBackOff on mofed podsKernel module conflictCheck thirdPartyRDMAModules, enable unloadThirdPartyRDMAModules
No VFs on nodeSriovNetworkNodePolicy not matchingVerify nodeSelector labels match worker nodes
RDMA not workingMissing RDMA device plugin or wrong resource nameCheck rdma-shared-dp pods, verify resource annotations
Phase 0 Helm chart download returns HTTP 401 or an image-pull-Secret errorThe configured Secret is missing from the operator namespace, unreadable by the kubeconfig, or has no compatible Docker auth entryVerify the Secret in networkOperator.namespace; for NGC, its .dockerconfigjson must contain nvcr.io credentials
l8k discover daemon pods stuck (ImagePullBackOff / Pending)Bad image tag, missing pull secret, or no feature.node.kubernetes.io/pci-15b3.present=true nodesRe-run with --keep-namespace then kubectl describe pod -n nvidia-k8s-launch-kit. Fix networkOperator.componentVersion / pass --image-pull-secrets / verify NFD is running.
l8k validate / deploy can't find Network Operator podsOperator namespace mismatchVerify --network-operator-namespace matches actual namespace (does NOT apply to l8k discover — it ignores the flag and uses its own nvidia-k8s-launch-kit namespace)
IPPool not allocatingNV-IPAM subnet exhausted or misconfiguredCheck ippools CR status, verify CIDR ranges
--for requires --node-selector--for was passed without --node-selectorAdd --node-selector key=val,…. The synthesized clusterConfig has no live worker-node list; the selector identifies target nodes at apply time.
--for and --discover-cluster-config are mutually exclusiveBoth flags passed simultaneouslyPick one: --for skips discovery, --discover-cluster-config runs it.
unknown preset "X"; available: …--for X doesn't match any directory under presets/Run l8k preset list and re-run with one of those names.
preset has no capabilities blockPreset YAML used by --for is missing capabilities.nodes.{sriov,rdma,ib}Add the block to the preset's topology.yaml. Discovery-time overlay does not require it; only --for does.
unknown field "productType" in YAMLHand-authored config still uses the old key nameRename productType: to gpuType: (the field was renamed).

For detailed triage workflow, read references/troubleshooting-guide.md.

sosreport Analysis

If the user has a pre-collected sosreport directory (from l8k sosreport or manual collection):

sosreport/
├── metadata/          # Cluster info, node list
├── crds/              # NicClusterPolicy, SriovNetworkNodePolicy, IPPool, etc.
├── operator/          # Network operator pod logs
├── nodes/             # Per-node device info
└── network/           # Interface config, routing tables

Triage Checklist

  1. Read metadata/diagnostic-summary.yaml for overview
  2. Check pod health in operator/pods.yaml
  3. Inspect CRDs in crds/ for status fields
  4. Read operator logs in operator/logs/ for errors
  5. Check per-node NIC state in nodes/<node>/

See Also

nvidia의 다른 스킬

compileiq-debug
nvidia
무언가 잘못되었을 때 사용: Search()가 멈추거나, 모든 평가가 INVALID_SCORE를 반환하거나, 점수가 개선되지 않거나, 모든 설정이 동일한 숫자를 반환하거나, ptxas 오류 등이 발생할 때
create-github-pr
nvidia
gh CLI를 사용하여 GitHub 풀 리퀘스트를 생성합니다. 사용자가 새 PR을 만들거나, 코드 리뷰를 제출하거나, 풀 리퀘스트를 열고자 할 때 사용합니다. 트리거 키워드 -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
다른 열린 이슈들을 스캔하여 주어진 PR이 함께 수정하거나 실수로 망가뜨릴 수 있는 이슈를 찾습니다. 인접 수정 기회와 모순 위험을 file:line…과 함께 출력합니다.
fhir-basics
nvidia
에이전트에게 FHIR R4 API의 작동 방식, 사용 가능한 리소스, 검색 매개변수를 사용한 쿼리 방법, 모든 응답 형식을 올바르게 파싱하는 방법을 가르칩니다…
compileiq-validate-result
nvidia
검색이 완료된 후, 속도 향상을 청구하거나 ACF를 발송하기 전에 사용합니다. dump_results CSV를 로드하고, 상위 K개 후보(단일 목표)를 추출합니다…
changelog-audit
nvidia
릴리스 전에 Warp CHANGELOG.md를 감사합니다: 누락된 항목 복구, 사용자 영향별 정렬, 항목 언어 다듬기, 줄 바꿈, (릴리스 브랜치 모드) 비교 업데이트…
maintain-dynamic-plugins
nvidia
NeMo Relay 동적 플러그인 로더, 매니페스트, Rust 네이티브 SDK, gRPC 워커 프로토콜, Python 워커 SDK, 문서, 테스트 및 릴리스 워크플로 커버리지를 유지 관리합니다.
dgx-diagnose
nvidia
일반적인 DGX Station GB300 문제 진단 — CUDA 충돌, 잘못된 GPU 타겟팅, vLLM/SGLang 컨테이너 버그, MIG 상태 문제, NVLink/Fabric Manager 오류,…