k8s-launch-kit-generate

작성자: nvidia

사용자가 k8s-launch-kit(l8k)을 사용하여 NVIDIA 네트워킹 배포를 위한 Kubernetes YAML 매니페스트를 생성하려 할 때 이 스킬을 사용하세요. 활성화 대상: 매니페스트…

npx skills add https://github.com/nvidia/k8s-launch-kit --skill k8s-launch-kit-generate

l8k: Manifest Generation

PREREQUISITE: Read ../k8s-launch-kit-shared/SKILL.md for install paths, global flags, and output modes.

This workflow uses the default host target; --target host is equivalent. The CLI snapshots generation arguments and runs the concrete Host generation operation through the target registry; command syntax and artifacts are unchanged.

Generate Kubernetes YAML manifests for NVIDIA networking from a cluster config and profile selection.

Usage

l8k generate --user-config <CONFIG> \
  --save-deployment-files <OUTPUT_DIR>

Configs produced by l8k discover already contain the resolved profile. Profile flags remain available as generation-time overrides. When generation uses a file-backed config, resolved defaults and CLI overrides are written back to that source file; embedded --for generation does not write a config.

Profile Selection Flags

FlagRequiredValuesDescription
--fabricAuto-defaultedethernet, infinibandNetwork fabric. Auto-defaults from the cluster's unanimous linkType when omitted (Unit 5 fabric probe); skipped+warned when groups disagree or any has unverified linkType.
--deployment-typeAuto-defaultedsriov, rdma_shared, host_deviceDeployment type. Auto-defaults to sriov.
--spectrum-xRA2.1, RA2.2, RA2.3Enable Spectrum-X profile by passing the SPC-X RA version. Implies ethernet fabric, sriov deployment, and multirail.
--multiplane-modeAuto-defaulted with --spectrum-xnone, swplb, hwplbAuto-defaults from GPU platform plus east-west PF deviceID: H100/H200/B200/GB200 → none; B300/GB300 → the GA swplb path. Platform cannot identify hwplb; select it explicitly. Unknown platforms fall back to NIC family.
--number-of-planesAuto-defaulted with --spectrum-x1, 2, 4Single-plane platforms → 1; B300/GB300 → 2. Pass 4 explicitly for quad-plane B300. An explicit none also implies 1, and an explicit 1 implies none.
--topology-schemeRequired with --spectrum-x2-tier, 3-tierSelects the Spectrum-X topology addressing scheme.
--ip-versionRequired with --spectrum-xipv4, ipv6Selects per-node IPv4 /31 or IPv6 /64 CIDRPool allocation.
--topology-fileRequired with --spectrum-xpathspcx-gen/reference-generator or contract-compliant NVIDIA AIR topology JSON. The format is detected from the JSON structure.
--multirailAuto-defaultedAuto-defaults to true. Explicit multirail: false in YAML and --multirail=false on the CLI are both preserved.
--save-deployment-filesYesOutput directory for generated YAMLs
--groupsdgx-b200-h100-nvl,pe-xe9680-h200Restrict output to the named source groups (comma-separated). Mutually exclusive with --gpu-type.
--gpu-typeNVIDIA-H200Restrict output to source groups whose gpuType matches (case-insensitive). Mutually exclusive with --groups.
--forpreset directory nameSkip discovery: synthesize clusterConfig from a topology preset. Requires --node-selector. List options with l8k preset list.
--node-selectorRequired with --forkey=val,key2=val2Identifies which nodes the synthesized clusterConfig targets at apply time.

*Not required when --spectrum-x is used.

Examples

# SR-IOV Ethernet RDMA (most common for GPU clusters)
l8k generate --user-config cluster-config.yaml \
  --fabric ethernet --deployment-type sriov \
  --save-deployment-files ./output

# Spectrum-X with hardware plane load balancing
l8k generate --user-config cluster-config.yaml \
  --spectrum-x RA2.2 --multiplane-mode hwplb --number-of-planes 4 \
  --topology-scheme 2-tier --ip-version ipv6 \
  --topology-file ./topology.json \
  --save-deployment-files ./output

# Host device RDMA
l8k generate --user-config cluster-config.yaml \
  --fabric ethernet --deployment-type host_device \
  --save-deployment-files ./output

# IPoIB RDMA shared (InfiniBand)
l8k generate --user-config cluster-config.yaml \
  --fabric infiniband --deployment-type rdma_shared \
  --save-deployment-files ./output

# Agent mode
l8k generate --user-config cluster-config.yaml \
  --fabric ethernet --deployment-type sriov \
  --save-deployment-files ./output \
  --output json 2>/dev/null

# Generate from a known server SKU (no cluster discovery required)
l8k preset list   # see available presets
l8k generate --user-config cluster-config.yaml \
  --for ThinkSystem-SR680a-V3 \
  --node-selector "nvidia.com/gpu.product=NVIDIA-H200" \
  --fabric ethernet --deployment-type sriov \
  --save-deployment-files ./output

Choosing l8k discover vs --for

  • l8k discover then l8k generate — default flow. Discovery learns the hardware and persists the resolved profile, so generation needs only the resulting cluster-config.yaml unless an override is desired.
  • l8k generate --for <preset> — skip discovery entirely when the SKU is already known and there is a preset for it. Useful for ahead-of-time generation (CI scaffolding, lab runbooks, demos), or when you don't have kubectl access yet. Requires --node-selector to identify the target nodes at apply time.

A preset used with --for must declare capabilities.nodes.{sriov,rdma,ib} in its topology.yaml. All bundled presets do.

Profile Quick Reference

ProfileFlagsUse Case
SR-IOV Ethernet RDMA--fabric ethernet --deployment-type sriovGPU clusters, ML training, HPC
Host Device RDMA--fabric ethernet --deployment-type host_deviceLegacy HPC, DPDK, full NIC access
MacVLAN RDMA Shared--fabric ethernet --deployment-type rdma_sharedMulti-tenant Ethernet environments
IPoIB RDMA Shared--fabric infiniband --deployment-type rdma_sharedInfiniBand shared workloads
SR-IOV InfiniBand--fabric infiniband --deployment-type sriovInfiniBand SR-IOV
Spectrum-X--spectrum-xAI cloud, multi-tenant GPU networking

For detailed profile selection guidance (NIC constraints, multiplane modes, when to use each), read references/profile-decision-tree.md.

Output

Generated YAMLs are written to the output directory under network-operator/. Each profile also emits a values.yaml (Helm values for the nvidia/network-operator chart) alongside the CR manifests:

When validation.gpuDirect.enabled is true, every generated example DaemonSet requests validation.gpuDirect.gpuResourceType only on its primary DOCA container. The request exposes the highest GPU<N> referenced by PF topology, the image comes from the selected release's validation.image, and networkOperator.imagePullSecrets is copied to the Pod spec. Do not inject these fields later at validation runtime.

output/
└── network-operator/
    ├── values.yaml                       # Phase 0 helm-install input for `l8k deploy`
    ├── 10-nicclusterpolicy.yaml
    ├── 11-nicnodepolicy-<group>.yaml
    ├── 20-ippool-<group>.yaml
    ├── 40-sriovnetworknodepolicy-<group>.yaml
    └── 50-sriovnetwork-<group>.yaml

values.yaml is rendered from the profile's 00-values.yaml template. --network-operator-release <MAJOR.MINOR> populates the chart repository URL and image tag from the embedded catalog. For Spectrum-X profiles, the same catalog entry supplies the independently versioned xPlane repository and tag rather than reusing the generic Network Operator component coordinates. To install or upgrade the chart alongside the CRs, pass --deploy (and --overwrite-existing when the release already exists with different values).

For Network Operator 26.1 and newer, the rendered values enable Maintenance Operator requestor mode. Profiles with DOCA/OFED enable operator.maintenanceOperator.useRequestor. Profiles that deploy the SR-IOV Operator enable both the Network Operator drain requestor and sriov-network-operator.operator.externalDrainer; the two SR-IOV switches are a coordinated handoff and must not be separated. The generated MaintenanceOperatorConfig gets the global limits from the config's maintenance section.

Before release 26.1, OFED uses maintenance.maxParallelUpgrades and the SR-IOV internal drainer uses maintenance.maxUnavailable through SriovNetworkPoolConfig. Starting with 26.1, the global Maintenance Operator limits are effective for both flows; the legacy OFED and SR-IOV pool limits do not control requestor-mode concurrency.

Reusing the Discovered Profile

l8k discover writes the final profile.fabric, profile.deployment, profile.multirail, and any enabled Spectrum-X settings to the config. Run generate without profile flags to reuse them; explicit generate flags still win when a one-off override is needed.

Common Mistakes

  • There is no --profile flag. Profiles are selected via --fabric + --deployment-type (or --spectrum-x). Do NOT invent flags.
  • The multiplane flag is --multiplane-mode, not --spcx-multiplane or --multiplane.

Tips

  • Default to SR-IOV Ethernet for new GPU cluster deployments unless told otherwise.
  • For Spectrum-X, GPU platform and NIC type determine the safe defaults, but B300/GB300 platform type does not distinguish SWPLB from HWPLB. l8k defaults to SWPLB; use an explicit HWPLB override when the site topology requires it. Read references/spectrum-x-modes.md.
  • Spectrum-X renders one NicConfigurationTemplate per source group and derives spec.nicSelector.nicType plus pciAddresses from that group's east-west PFs. The intersection prevents same-device-ID north-south DPUs from matching. Do not ask users to configure spectrumX.nicType; missing or mixed east-west device IDs and missing east-west PCI addresses are generation errors.
  • NVIDIA AIR topology support requires the documented one-based node/interface naming contract (su<S>, h<H>, leaf-p<P>, r<R>, rail<R>p<P>, and pod<D> for 3-tier). See docs/user/spectrum-x.md in the l8k repository.
  • Spectrum-X CIDRPool allocation matches clusterConfig.workerNodes to topology host endpoint node values exactly and case-sensitively. A zero-match error usually means the wrong topology file, a case difference, or a short-name/FQDN difference. Partial-pool errors report the worker's available rail/plane coverage; check host attributes.rail and, for swplb, leaf attributes.plane.
  • RA2.2 and RA2.3 v1alpha2 SpectrumXRailPoolConfig output intentionally omits the removed spec.withBCM field; current CRDs reject it during strict decoding.
  • Group identifiers produced by discovery omit complete NVIDIA segments and shorten common machine segments (ThinkSystemts, PowerEdgepe). They are bounded to 30 bytes with balanced machine/GPU prefixes and a 6-character deterministic hash suffix. The machine label uses the same value; the GPU label still retains its discovered value such as NVIDIA-H200. Use the exact persisted identifier from cluster-config.yaml with --groups; do not reconstruct it from long machineType and gpuType strings.
  • Use --groups <a,b,...> (case-sensitive identifier list) or --gpu-type <X> (case-insensitive) to scope a generate to a subset of source groups in heterogeneous clusters. Mutually exclusive. Empty match is a validation error. Strict-subset filters split per-source rendering: NodePolicies emit one CR per source (each with its own machine-label nodeSelector but a shared bucket-level resourceName); IPPool/example DaemonSet emit one CR per bucket with an In list of source machine labels.

[!CAUTION] Generation does not apply anything to the cluster. Use --deploy or k8s-launch-kit-deploy to apply.

See Also

nvidia의 다른 스킬

compileiq-debug
nvidia
무언가 잘못되었을 때 사용: Search()가 멈추거나, 모든 평가가 INVALID_SCORE를 반환하거나, 점수가 개선되지 않거나, 모든 설정이 동일한 숫자를 반환하거나, ptxas 오류 등이 발생할 때
create-github-pr
nvidia
gh CLI를 사용하여 GitHub 풀 리퀘스트를 생성합니다. 사용자가 새 PR을 만들거나, 코드 리뷰를 제출하거나, 풀 리퀘스트를 열고자 할 때 사용합니다. 트리거 키워드 -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
다른 열린 이슈들을 스캔하여 주어진 PR이 함께 수정하거나 실수로 망가뜨릴 수 있는 이슈를 찾습니다. 인접 수정 기회와 모순 위험을 file:line…과 함께 출력합니다.
fhir-basics
nvidia
에이전트에게 FHIR R4 API의 작동 방식, 사용 가능한 리소스, 검색 매개변수를 사용한 쿼리 방법, 모든 응답 형식을 올바르게 파싱하는 방법을 가르칩니다…
compileiq-validate-result
nvidia
검색이 완료된 후, 속도 향상을 청구하거나 ACF를 발송하기 전에 사용합니다. dump_results CSV를 로드하고, 상위 K개 후보(단일 목표)를 추출합니다…
changelog-audit
nvidia
릴리스 전에 Warp CHANGELOG.md를 감사합니다: 누락된 항목 복구, 사용자 영향별 정렬, 항목 언어 다듬기, 줄 바꿈, (릴리스 브랜치 모드) 비교 업데이트…
maintain-dynamic-plugins
nvidia
NeMo Relay 동적 플러그인 로더, 매니페스트, Rust 네이티브 SDK, gRPC 워커 프로토콜, Python 워커 SDK, 문서, 테스트 및 릴리스 워크플로 커버리지를 유지 관리합니다.
dgx-diagnose
nvidia
일반적인 DGX Station GB300 문제 진단 — CUDA 충돌, 잘못된 GPU 타겟팅, vLLM/SGLang 컨테이너 버그, MIG 상태 문제, NVLink/Fabric Manager 오류,…