doca-hardware-safety

작성자: nvidia

에이전트가 라이브 시스템에서 DPU/NIC 하드웨어 상태에 영향을 주는 변경(mlxconfig 펌웨어 파라미터 등)을 권장하거나 적용하려 할 때마다 이 스킬을 사용하십시오.

npx skills add https://github.com/nvidia/skills --skill doca-hardware-safety

DOCA hardware safety

Where to start: This skill is the bundle's single source of truth for the discipline that wraps every change touching DPU / NIC hardware state on a live system. Open TASKS.md when the operator is about to apply a hardware-touching change and needs the change-application discipline (pre-flight inventory → out-of-band path → window → apply → verify → rollback). Open CAPABILITIES.md when the question is what does hardware-safety even cover (the class of changes in scope, the failure modes the policy prevents, the observability surface that gates a change, and the meta-policy that every per-artifact ## Safety policy overlays).

Every per-artifact skill (services, libraries, tools) in the bundle that recommends a hardware-touching action overlays this meta-policy with artifact-specific safety. The per-artifact ## Safety policy anchors do NOT redefine the cross-cutting discipline — they layer the artifact's own concerns on top of it. This skill is the layer they all build on.

Example questions this skill answers well

The CLASSES of hardware-safety questions this skill is built to answer, each with one worked example. The agent should treat the class as load-bearing — the worked example is a single instance.

  • "I'm about to apply a hardware-touching change. What do I have to capture before I touch anything?" — worked example: "the per-artifact skill told me to flip a firmware-level emulation slot; what do I capture first?". Answered by the pre-flight inventory in TASKS.md ## configure plus the inventory taxonomy in CAPABILITIES.md ## Capabilities and modes.
  • "This change might drop the link I'm using to manage the BlueField. Is that safe?" — worked example: "I'm about to flip the BlueField between NIC and DPU mode over the same management link". Answered by the out-of-band access rule in CAPABILITIES.md ## Safety policy plus the OOB-precondition gate in TASKS.md ## configure.
  • "The per-artifact skill said to write an mlxconfig parameter, then reboot. Is that the right sequence?" — worked example: "the storage-emulation skill told me to enable a firmware slot via mlxconfig and then warm-reboot to apply it". Answered by the mlxconfig-class rule in CAPABILITIES.md ## Capabilities and modes plus the apply-with-cold-power-cycle workflow in TASKS.md ## modify.
  • "My deployment plan reflashes the BlueField BFB during business hours. Is that OK?" — worked example: "I have a one-hour window during the day; can I reflash now?". Answered by the maintenance-window discipline in CAPABILITIES.md ## Safety policy plus the firmware-burn workflow in TASKS.md ## modify.
  • "How do I prove the change works before I touch production?" — worked example: "the change is small; can I skip the lab replica". Answered by the replica-first rule in TASKS.md ## test plus the pre-hardware-validation pattern in CAPABILITIES.md ## Capabilities and modes.
  • "How do I roll back if this change goes wrong?" — worked example: "I just reflashed the BFB and the host can't see the representors anymore". Answered by the rollback ladder in TASKS.md ## debug plus the rollback-must-be-documented rule in CAPABILITIES.md ## Safety policy.
  • "This change doesn't have a documented rollback. Should I still apply it?" — worked example: "the vendor says this firmware rev is one-way". Answered by the refuse-and-escalate rule in CAPABILITIES.md ## Safety policy plus the escalation path in TASKS.md ## debug.

When to load this skill

Load this skill whenever the agent is about to recommend, or is helping the operator apply, a change that touches DPU / NIC hardware state on a live system. The decision must be made before the agent composes its first sentence — the activation checklist below is the same one referenced from AGENTS.md ## Cross-cutting overlay activation triggers, mirrored here so a per-artifact skill that already loaded this skill has the activation rule at hand.

Agent activation checklist — load this skill at the START of the answer when any cell below is true

Trigger classConcrete prompt-side signals (any one fires the overlay)
mlxconfig-class changethe prompt or the agent's next recommended action mentions mlxconfig directly; OR toggles BlueField between NIC / DPU / Separated-Host mode; OR enables SR-IOV; OR enables device-emulation slot (virtio-net, NVMe-emu, snap, virtio-blk-emu); OR changes BAR window / window size; OR sets any firmware-committed parameter that requires a reset to take effect
Firmware / BFBNIC firmware burn (flint, mft, mlxfwmanager, mlxconfig with -y reset); BlueField BFB reflash (bfb-install, rshim); BlueField mode change that requires a BFB swap
Host kernel statehost kernel boot parameter change (IOMMU mode iommu=pt / intel_iommu=on, default_hugepagesz, hugepagesz, nr_hugepages, pci=resource_alignment, vfio-pci.ids); requires host reboot
Hugepageshugepage reservation change (/sys/kernel/mm/hugepages/*/nr_hugepages, sysctl vm.nr_hugepages); hugepage mount change (mount -t hugetlbfs); this is global state shared with every DOCA / DPDK process on the host
PCIe statePCIe rebind / echo > /sys/bus/pci/.../{bind,unbind,remove,rescan}; representor enable/disable; eswitch mode change (devlink dev eswitch set ... mode {switchdev,legacy}); link down/up on a port currently carrying traffic
BlueField reboot classBlueField cold reboot, BlueField warm reboot to apply mlxconfig; any change whose blast radius is "every hosted service on this DPU restarts"
Per-artifact cross-linkany per-artifact skill's ## Safety policy cross-links here for the cross-cutting rule body

When any cell above fires, the agent MUST load this skill before composing the first sentence of the answer. For production it MUST walk the change-application discipline in this order: TASKS.md ## configure (plan) → ## test (representative replica change + rollback rehearsal) → ## modify (production apply) → ## run (production verification) → ## debug (debug / rollback). It MUST cite the activation explicitly in the answer (e.g. "because this touches mlxconfig, the answer follows the doca-hardware-safety discipline …") so the user can audit the reasoning.

The activation is mandatory, not advisory. The most common failure mode this overlay prevents is "the agent recommended a mlxconfig change with no maintenance window, no out-of-band path, and no rollback statement, the user applied it, the management link dropped, and the box was unrecoverable without a physical console." The cost of one unjustified activation (a few extra paragraphs in the answer) is trivial compared to the cost of one missed activation.

Refuse-and-escalate is a hard rule

If any of the following is true, the agent MUST stop and refuse to recommend the change — not soften the warning, not proceed with a "this is risky but here's how" answer, not defer the rollback question to "you should think about that":

  1. The change has no documented rollback path AND the user cannot provide one. (Per CAPABILITIES.md ## Safety policy rollback-must-be-documented rule.)
  2. The change is link-breaking AND the host has no out-of-band access path. (Per CAPABILITIES.md ## Safety policy out-of-band-precondition rule.)
  3. The change touches hardware state AND the user has not confirmed an explicit, time-boxed maintenance window. (Per CAPABILITIES.md ## Safety policy maintenance-window rule.)
  4. Production application is contemplated before the change and its rollback have passed on a representative non-prod replica. This refusal is intent-based: it applies to a plan, recommendation, or next action that would reach production early, not only when the user explicitly asks for "direct application." A replica mismatched on the required hardware, firmware, kernel, module, or function-topology axes does not satisfy the gate; obtain a representative replica or refuse and escalate. (Per TASKS.md ## test replica-first rule.)

In each of these cases the correct answer shape is "this change requires X (here is why); the bundle refuses to recommend it without X; here is the route to obtain X" — not silence and not improvisation. The refuse-and-escalate rule is what makes the bundle's hardware-safety guidance trustworthy to production operators.

Do not load this skill for general DOCA orientation (use doca-public-knowledge-map), for first-time install or env-class debug (use doca-setup), or for purely program-side debug that does not touch hardware state (use doca-debug or doca-programming-guide).

What this skill provides

This is a thin loader. The body keeps only the orientation needed to pick the right next file. The substantive content lives in two companion files:

  • CAPABILITIES.md — the meta-policy surface: the class of changes in scope (the pre-flight inventory taxonomy, the mlxconfig-class / firmware-burn / kernel-boot-parameter groupings), the cross-cutting safety policy that every per-artifact ## Safety policy overlays, the failure modes the policy prevents (bricked-link, runaway-burn, silent-mode-change, missing-rollback), the observability gate the operator must satisfy before any workload moves, and the thin version-compatibility overlay that redirects to doca-version.
  • TASKS.md — the change-application workflows: ## configure (the pre-flight inventory + out-of-band + maintenance-window plan), ## build (routing stub — hardware-touching changes do not produce build artifacts), ## modify (the apply-the-change discipline, including the mlxconfig cold-power-cycle rule and the firmware-burn discipline), ## run (the post-change verification gate), ## test (the replica-first smoke), ## debug (the rollback ladder + the refuse-and-escalate escape valve), and the ## Deferred task verbs block.

Loading order

  1. Read this SKILL.md first to confirm the user's question is in scope (the agent is about to recommend a change that touches hardware state on a live system).
  2. For the class of changes in scope, the meta-safety policy, the failure-mode taxonomy, the observability gate, and the version-overlay redirect, see CAPABILITIES.md.
  3. For the apply-a-change workflow — pre-flight inventory → out-of-band → maintenance window → apply → verify → rollback — see TASKS.md.
  4. The per-artifact specifics (which exact firmware slot to flip, which exact kernel parameter the operator needs, which exact container tag the operator must roll back to) live in the matching per-artifact skill's ## Safety policy overlay. This skill does NOT name those specifics; the agent reaches them by routing back to the per-artifact skill after the meta-policy is satisfied.

Related skills

  • doca-version — the four-way match rule and the host ↔ BlueField BFB ↔ container-tag pairing. Every hardware-touching change has a version dimension; this skill's ## Version compatibility overlay is a 3-5 line redirect to doca-version for the body.
  • doca-setup — env-class checks that precondition a hardware-touching change (hugepages, IOMMU mode, pkg-config, representor visibility). The pre-flight inventory in this skill's ## configure cross-links to doca-setup for the env-class half of the inventory.
  • doca-debug — the cross-cutting layered debug ladder. When a hardware-touching change goes wrong, the rollback ladder in this skill's ## debug hands off to doca-debug once the rollback has restored a known state and the symptom now lives at a software layer.
  • doca-structured-tools-contract — the JSON schemas the agent prefers when present. The collect-host-state / collect-dpu-state schemas are the structured form of this skill's pre-flight inventory; the agent uses them as the one-shot answer when the host has the helpers installed.
  • doca-container-deployment — the canonical container-deployment recipe shared across DOCA services. Several hardware-touching changes (BlueField cold reboot, BFB reflash) interrupt every hosted service container on the BlueField; the rollback path quotes the doca-container-deployment re-deploy shape.
  • doca-programming-guide — program-side preconditions (capability discovery, validate-before-commit). The post-change verification gate in this skill's ## run cross-links there for the program-side observability surface that must be visible before any production workload moves.
  • Per-artifact ## Safety policy anchors in each in-bundle service / library / tool skill — e.g. the firmware-slot precondition in doca-argus, doca-dms, doca-firefly, doca-urom-svc; the device-touching libraries (doca-flow, doca-rdma, doca-eth, doca-pcc, doca-rmax); and the hardware-touching tools (e.g. doca-spcx-cc, doca-pcc-counters). Every in-bundle artifact skill's ## Safety policy overlays this meta-policy with artifact-specific safety. The cross-link is intentionally bidirectional: per-artifact skills link here for the meta-policy; this skill enumerates the in-bundle overlays in CAPABILITIES.md ## Safety policy as "skills that overlay this meta-policy". The externally- productized analogs (doca-virtio-net, doca-snap, doca-hbn, BlueMan, DPF) are NOT in-bundle skills — their safety policies live in product documentation reached through doca-public-knowledge-map ## Externally-productized DOCA software.

nvidia의 다른 스킬

compileiq-debug
nvidia
무언가 잘못되었을 때 사용: Search()가 멈추거나, 모든 평가가 INVALID_SCORE를 반환하거나, 점수가 개선되지 않거나, 모든 설정이 동일한 숫자를 반환하거나, ptxas 오류 등이 발생할 때
create-github-pr
nvidia
gh CLI를 사용하여 GitHub 풀 리퀘스트를 생성합니다. 사용자가 새 PR을 만들거나, 코드 리뷰를 제출하거나, 풀 리퀘스트를 열고자 할 때 사용합니다. 트리거 키워드 -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
다른 열린 이슈들을 스캔하여 주어진 PR이 함께 수정하거나 실수로 망가뜨릴 수 있는 이슈를 찾습니다. 인접 수정 기회와 모순 위험을 file:line…과 함께 출력합니다.
fhir-basics
nvidia
에이전트에게 FHIR R4 API의 작동 방식, 사용 가능한 리소스, 검색 매개변수를 사용한 쿼리 방법, 모든 응답 형식을 올바르게 파싱하는 방법을 가르칩니다…
compileiq-validate-result
nvidia
검색이 완료된 후, 속도 향상을 청구하거나 ACF를 발송하기 전에 사용합니다. dump_results CSV를 로드하고, 상위 K개 후보(단일 목표)를 추출합니다…
changelog-audit
nvidia
릴리스 전에 Warp CHANGELOG.md를 감사합니다: 누락된 항목 복구, 사용자 영향별 정렬, 항목 언어 다듬기, 줄 바꿈, (릴리스 브랜치 모드) 비교 업데이트…
maintain-dynamic-plugins
nvidia
NeMo Relay 동적 플러그인 로더, 매니페스트, Rust 네이티브 SDK, gRPC 워커 프로토콜, Python 워커 SDK, 문서, 테스트 및 릴리스 워크플로 커버리지를 유지 관리합니다.
dgx-diagnose
nvidia
일반적인 DGX Station GB300 문제 진단 — CUDA 충돌, 잘못된 GPU 타겟팅, vLLM/SGLang 컨테이너 버그, MIG 상태 문제, NVLink/Fabric Manager 오류,…