k8s-launch-kit-shared

작성자: nvidia

k8s-launch-kit (l8k) CLI: 바이너리 위치, 전역 플래그, 출력 형식, 종료 코드 및 오류 처리를 위한 공유 패턴입니다. 사용하기 전에 이 내용을 읽어보세요…

npx skills add https://github.com/nvidia/k8s-launch-kit --skill k8s-launch-kit-shared

l8k — Shared Reference

Installation

# Build and install (production — copies binary + profiles)
make build && sudo scripts/install.sh

# Development (symlinks into source tree)
make build && scripts/install.sh --dev-env

Install Paths

PathContents
/usr/local/bin/l8kCLI binary (on PATH)
/usr/local/share/l8k/profiles/Go template profiles
/usr/local/share/l8k/presets/Topology presets (per (machineType, gpuType) directories)
/usr/local/share/l8k/l8k-config.yamlDefault config

After installation, l8k is available system-wide.

Binary Discovery (for AI agents)

CRITICAL — follow these steps exactly:

  1. Run which l8k in a single Bash call. If it returns a path, use it. Done.
  2. If which l8k fails (exit code 1), tell the user l8k is not installed. Point them to: make build && sudo scripts/install.sh. Stop — do not proceed.
  3. NEVER do any of the following to find l8k:
    • find or ls commands
    • Spawn a subagent to locate the binary
    • Check multiple paths or directories
    • Explore the repo contents

Available Commands

CommandDescription
l8k discoverDiscover cluster hardware and produce cluster-config.yaml
l8k generateGenerate Kubernetes YAML manifests from config + profile (use --for <preset> to skip cluster discovery for known SKUs)
l8k deployApply generated manifests and install or upgrade the Network Operator Helm release
l8k cleanDelete Network Operator custom resources and optionally uninstall its Helm release
l8k validateVerify the Helm release, manifests, component versions, and connectivity
l8k preset listList bundled topology presets (directory + machineType + gpuType)
l8k preset updateDownload latest topology presets from GitHub
l8k sosreportCollect diagnostic dump from cluster
l8k schemaList all capabilities as JSON (profiles, flags, exit codes)
l8k versionPrint version information

The root command l8k --discover-cluster-config ... still works for backward-compatible full-pipeline usage.

Target selection

discover, generate, deploy, validate, and the root pipeline default to the host target. Existing invocations must omit --target unless the user explicitly requests it; --target host is an equivalent explicit form.

The dpf target name is reserved but its phases are not implemented in this build. Do not attempt DPF provisioning. Use l8k schema and inspect targets[].phases before selecting any non-host target. Host-only flags such as --fabric, --kubeconfig, and the Network Operator/Spectrum-X flags are rejected when explicitly combined with another target.

Internally, each lifecycle command snapshots the explicitly supplied Host arguments, binds a phase-specific adapter through the target registry, and runs the resulting operation exactly once with the command context. Do not bypass this route when extending or automating a lifecycle phase. The canonical component and data-flow diagrams live in docs/architecture/overview.md; update them with any change to lifecycle, package ownership, artifacts, or external integration boundaries.

Global Flags

FlagDescription
--target <NAME>Lifecycle target: host (default); dpf is currently unavailable
--kubeconfig <PATH>Path to kubeconfig file (optional — falls back to $KUBECONFIG env var)
--user-config <PATH>Path to user-supplied l8k-config.yaml
--output <FORMAT>Output format: text (default), json
--yes / -yAuto-confirm all prompts (root command only — not available on subcommands; --output json auto-confirms)
--quiet / -qSuppress informational output (errors still shown)
--log-level <LEVEL>Enable logs at trace, debug, info, warn, or error. Debug shows structured progress; trace also shows bounded command output.
--network-operator-namespace <NS>Override network operator namespace (default: nvidia-network-operator). No-op for l8k discover — discover always bootstraps into nvidia-k8s-launch-kit; the flag still applies to l8k generate / l8k deploy / l8k clean / l8k validate.
--network-namespaces <NS,...>Comma-separated namespaces for the secondary-network CRs + example test DaemonSets; one copy rendered per namespace (shared resources like IPPools/NodePolicies are NOT duplicated). Default: default
--node-selector <LABELS>Restrict to nodes matching labels (comma-separated, ANDed)
--image-pull-secrets <NAMES>Image pull secret names for Network Operator components and authenticated Helm chart downloads (comma-separated)

l8k discover and l8k generate both accept the profile flags --fabric, --deployment-type, --multirail, --spectrum-x, --multiplane-mode, and --number-of-planes. Discovery persists the resolved values; later generation reuses them unless another explicit CLI override is supplied.

Agent / JSON Mode

RULE: AI agents MUST always use --output json 2>/dev/null when calling any l8k subcommand. Never use text mode — it produces unstructured output with spinners and ANSI codes that is hard to parse.

Do NOT use --yes with subcommands — it only exists on the root command and will cause "unknown flag" errors. --output json already auto-confirms all prompts (no interactive input needed).

# Correct
l8k discover --output json 2>/dev/null
l8k generate --output json 2>/dev/null

# WRONG — --yes is not a valid flag on subcommands
l8k discover --output json --yes 2>/dev/null
  • stdout: Exactly one JSON object (JSONResult) at completion
  • stderr: Human-readable log lines
  • Pipe with jq for downstream processing: l8k discover ... --output json 2>/dev/null | jq .success

Exit Codes

CodeMeaning
0Success
1General/unknown error
2Validation error — bad flags, missing required arguments, or invalid config
3Cluster error — kubeconfig invalid, API unreachable, missing CRDs, no NVIDIA NICs
4Deployment error — apply failed
5Partial success — discovery ok but deploy failed

Structured Error Output (JSON mode)

{
  "success": false,
  "phase": "discover",
  "deployed": false,
  "error": {
    "code": "CLUSTER_ERROR",
    "message": "cluster discovery failed",
    "category": "cluster",
    "transient": true,
    "suggestion": "Check that kubeconfig is valid and the cluster is reachable"
  },
  "messages": [...]
}

Error categories: validation, cluster, deployment. The transient field hints whether retrying might help.

Schema Discovery

# List all l8k capabilities as JSON
l8k schema

Use l8k schema to programmatically discover available profiles, fabrics, deployment types, flags, exit codes, and output formats.

Security Rules

  • Confirm with the user before executing --deploy on a production cluster
  • Prefer --dry-run for destructive operations
  • Never expose kubeconfig contents in output

Network Operator Namespace Resolution

Applies to l8k generate / l8k deploy / l8k validate only. l8k discover manages its own private namespace (nvidia-k8s-launch-kit) and ignores this flag.

Both nvidia-network-operator and network-operator are common default namespaces for an existing Network Operator install. If l8k generate / l8k deploy / l8k validate can't find Network Operator resources, retry with --network-operator-namespace <correct-namespace>.

nvidia의 다른 스킬

compileiq-debug
nvidia
무언가 잘못되었을 때 사용: Search()가 멈추거나, 모든 평가가 INVALID_SCORE를 반환하거나, 점수가 개선되지 않거나, 모든 설정이 동일한 숫자를 반환하거나, ptxas 오류 등이 발생할 때
create-github-pr
nvidia
gh CLI를 사용하여 GitHub 풀 리퀘스트를 생성합니다. 사용자가 새 PR을 만들거나, 코드 리뷰를 제출하거나, 풀 리퀘스트를 열고자 할 때 사용합니다. 트리거 키워드 -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
다른 열린 이슈들을 스캔하여 주어진 PR이 함께 수정하거나 실수로 망가뜨릴 수 있는 이슈를 찾습니다. 인접 수정 기회와 모순 위험을 file:line…과 함께 출력합니다.
fhir-basics
nvidia
에이전트에게 FHIR R4 API의 작동 방식, 사용 가능한 리소스, 검색 매개변수를 사용한 쿼리 방법, 모든 응답 형식을 올바르게 파싱하는 방법을 가르칩니다…
compileiq-validate-result
nvidia
검색이 완료된 후, 속도 향상을 청구하거나 ACF를 발송하기 전에 사용합니다. dump_results CSV를 로드하고, 상위 K개 후보(단일 목표)를 추출합니다…
changelog-audit
nvidia
릴리스 전에 Warp CHANGELOG.md를 감사합니다: 누락된 항목 복구, 사용자 영향별 정렬, 항목 언어 다듬기, 줄 바꿈, (릴리스 브랜치 모드) 비교 업데이트…
maintain-dynamic-plugins
nvidia
NeMo Relay 동적 플러그인 로더, 매니페스트, Rust 네이티브 SDK, gRPC 워커 프로토콜, Python 워커 SDK, 문서, 테스트 및 릴리스 워크플로 커버리지를 유지 관리합니다.
dgx-diagnose
nvidia
일반적인 DGX Station GB300 문제 진단 — CUDA 충돌, 잘못된 GPU 타겟팅, vLLM/SGLang 컨테이너 버그, MIG 상태 문제, NVLink/Fabric Manager 오류,…