doca-bench-extension

작성자: nvidia

운영자가 사용자 정의 doca-bench 플러그인(버전이 지정된 공유 라이브러리)을 작성, 빌드, 로드 또는 디버깅할 때 이 스킬을 사용하십시오…

npx skills add https://github.com/nvidia/skills --skill doca-bench-extension

DOCA Bench Extension

Where to start: This is a tool skill for the extension / plug-in framework that augments doca-bench — NOT a workload-shape skill on its own. Open TASKS.md and start at ## configure to commit to the three-axis decision (workload class is genuinely outside doca-bench's built-in modes × extension API surface fits × parent-tool co-load is acceptable), then ## build for how a custom extension is compiled and laid out, then ## run for how doca-bench discovers and invokes the extension, then ## test for the smoke-before-bulk loop the agent applies to every new extension. Open CAPABILITIES.md when the question is what an extension can do that built-in doca-bench modes cannot, what the extension API surface looks like in broad strokes (the DOCA_EXPERIMENTAL C entry points the shipped reference exposes), how the build / registration / discovery flow works, or how the extension's lifetime is bounded by the parent doca-bench invocation. If doca-bench itself is the question, route to doca-bench. If the question is "which built-in doca-bench mode do I pick?", that is also doca-bench — extensions are the exit ramp for workloads built-in modes do not cover.

Example questions this skill answers well

  • "My workload class is <X> — does doca-bench measure it natively, or do I need an extension?" — the extension-vs-built-in decision question. The agent walks the user back to doca-bench's built-in mode inventory FIRST and only routes to the extension framework when no built-in mode applies.
  • "I want to benchmark a CUDA / GPU-side workload that drives DOCA GPUNetIO RX and TX queues. Where do I start? Is there a reference extension I can copy?" — the agent surfaces the shipped doca_bench_cuda extension under /opt/mellanox/doca/tools/bench_extension/doca_bench_cuda/ as the reference exemplar and walks the operator through its API surface and build shape.
  • "How does doca-bench actually discover and load my custom extension at runtime? Is it a versioned shared library? What does my entry-point need to look like?" — the build / registration / discovery flow question. The agent walks the Meson-built shared library shape, the versioning, and the parent-tool's runtime discovery path (which the agent does NOT invent from memory — the shipped extension's meson.build and the public DOCA Bench documentation on docs.nvidia.com are the source of truth).
  • "The API headers I have are marked DOCA_EXPERIMENTAL. What does that mean for my extension's stability across DOCA releases? Am I going to have to rebuild it every release?" — the experimental-surface and version compatibility question.
  • "Once I build my extension, what is the cheapest possible smoke I can run before pointing my real workload at it? How do I know doca-bench actually loaded it, called into it, and that the call returned the data the parent tool expected?" — the smoke-before-bulk question.
  • "My custom extension builds, but doca-bench says it cannot find / load / call it. Where do I look first?" — the layered-debug question that distinguishes build-failures, load-failures, registration-mismatches, and runtime-call-failures.

Audience

Experienced AI agents and platform / performance engineers who already use doca-bench for the built-in workload modes and now have a workload class that the built-in modes do not cover. Readers are expected to be comfortable with native build systems (Meson, in this codebase), shared-library packaging on Linux, and the DOCA_EXPERIMENTAL API stability contract. If the user asks about GPU-side benchmarking via the shipped doca_bench_cuda reference extension, the reader is also expected to be familiar with DOCA GPUNetIO and CUDA toolchain basics — those domains live in their own skills, not here.

This skill is NOT for:

  • operators who can express their workload with one of doca-bench's built-in modes — that is doca-bench;
  • operators who want to benchmark a different DOCA primitive (Flow, Comch, RMAX) via that primitive's own measurement tool — route to that tool;
  • contributors authoring or modifying the in-tree extensions themselves (this skill is for external operators consuming the framework, not for internal DOCA contributors).

Language scope

A doca-bench extension surfaces as:

  1. A versioned shared library on Linux (.so with soversion matching the DOCA release), built via the doca-bench-extension Meson rules in the shipped /opt/mellanox/doca/tools/bench_extension/meson.build and the per-extension subdirectory (the reference exemplar is doca_bench_cuda/).
  2. A small set of DOCA_EXPERIMENTAL-marked C entry points that the parent doca-bench invokes — i.e. the API surface declared in the extension's header file. The shipped doca_bench_cuda/doca_bench_cuda.h is the reference for what that surface shape looks like in practice (*_init, *_device_query, *_device_synchronize, and per-workload kernel-start entry points such as *_start_nop_kernel, *_start_eth_recv_kernel, *_start_eth_send_kernel, *_start_eth_bidir_kernel).
  3. A set of per-workload settings structs that the parent passes through (e.g. the reference exemplar's doca_bench_cuda_kernel_settings, doca_bench_cuda_eth_rx_kernel_settings, doca_bench_cuda_eth_tx_kernel_settings, doca_bench_cuda_eth_bidir_kernel_settings carry block counts, threads-per-block, RX / TX queues, buffer address / mkey / size, a stop flag, and a stats pointer).

The skill itself is Markdown. The user's extension source is whatever language the workload requires (C / C++ / CUDA in the reference case). The agent does NOT prescribe a language beyond what the shipped reference demonstrates.

When to load this skill

Load doca-bench-extension when ANY of the following is true:

  • the user explicitly mentions doca-bench-extension, the doca_bench_cuda reference extension, the doca_bench_cuda_impl shared library, or any of the DOCA_EXPERIMENTAL extension entry points;
  • the user has confirmed (via doca-bench TASKS.md ## configure) that none of doca-bench's built-in workload modes measures the class they want, and an extension is the exit ramp;
  • the user wants to copy / extend the shipped doca_bench_cuda reference into a custom GPU-side workload extension;
  • the user is debugging why doca-bench cannot find / load / call a custom extension they built.

Co-load this skill with:

  • doca-bench (the parent tool — ALWAYS co-loaded; extensions only have value as plug-ins into doca-bench);
  • doca-version (the DOCA_EXPERIMENTAL surface is versioned with DOCA; the extension's soversion is the DOCA soversion; the four-way version match applies);
  • doca-gpunetio when the extension is GPU-side and uses GPUNetIO RX / TX queues like the reference exemplar (route the GPUNetIO semantics there, not here);
  • doca-debug and doca-setup for the env-side debug ladder (driver, firmware, CUDA toolkit, dynamic linker).

Do NOT load this skill when the user's workload fits a doca-bench built-in mode — extensions add cost (build toolchain, version churn, the experimental-surface contract); the built-in modes are always the first answer to try.

What this skill provides

Three companion files in this directory, each owning a different question shape:

  • SKILL.md — this file. Audience, scope, loading order, related skills. Routes everything else.
  • CAPABILITIES.mdwhat an extension can do that the built-in modes cannot, what the API surface looks like in broad strokes, how the build / registration / discovery flow works, what versions it ships in (including the DOCA_EXPERIMENTAL-stability overlay on top of doca-version), the layered error taxonomy, observability, and the safety policy overlay.
  • TASKS.md — the procedural verbs (configure, build, run, test, debug, etc.) plus a doca-bench-extension-specific command appendix and the agent-side use workflow that consumes the captured extension run.

The combined skill teaches an AI agent to drive the extension-author-and-wire-in class of doca-bench questions: confirm an extension is needed at all; locate the shipped reference exemplar (/opt/mellanox/doca/tools/bench_extension/doca_bench_cuda/); copy its build + API surface shape; build a versioned shared library that matches the DOCA release; smoke that the parent doca-bench actually loads it; diagnose layered failures when it does not.

What this skill deliberately does not ship

  • Inventory of doca-bench's built-in workload modes. That belongs to doca-bench. This skill is the exit ramp for what the built-in modes do not cover; it does not duplicate the parent's mode inventory.
  • Invented DOCA_EXPERIMENTAL entry-point names beyond what the shipped reference declares. The shipped doca_bench_cuda/doca_bench_cuda.h on the user's install is the reference for what the surface shape looks like; the agent does not assert other extensions exist with specific signatures.
  • A canonical "right" extension layout. The shipped doca_bench_cuda reference IS the canonical layout; rewriting it here would drift from the source of truth. The agent points the operator at the shipped tree and walks the operator through adapting it.
  • A documented runtime discovery mechanism the agent invents. The exact mechanism doca-bench uses to locate and load extensions (search path, naming convention, registration call) lives in the public DOCA Bench documentation on docs.nvidia.com and the installed doca-bench binary. The agent points the operator there rather than asserting a mechanism from memory.
  • DOCA GPUNetIO programming details. When the extension is GPU-side (as the reference exemplar is), the GPUNetIO RX / TX queue semantics live in doca-gpunetio; this skill cross-links rather than duplicates.
  • CUDA toolchain installation guidance. Route to the public NVIDIA CUDA Toolkit documentation on docs.nvidia.com; this skill does not duplicate it.
  • Library-internal doca-bench invocation details unrelated to extensions. The parent's CLI flags, pipeline shapes, and built-in workload classes belong to doca-bench.

Loading order

When a doca-bench-extension question arrives:

  1. Confirm DOCA is installed AND doca-bench is reachable on the user's install — if not, route to doca-setup;
  2. Confirm none of doca-bench's built-in modes covers the workload class — if any of them does, route back to doca-bench TASKS.md ## configure and stop. Extensions are the exit ramp, not the first answer;
  3. Read CAPABILITIES.md to commit to the three-axis decision and walk the reference exemplar's API surface shape;
  4. Read TASKS.md and walk ## configure → ## build → ## run → ## test → ## debug in that order; do NOT start with ## run without the build precondition step.

Related skills

Cross-link conventions follow the bundle's relative path contract from tools/<X>/:

  • doca-bench — the parent tool. ALWAYS co-loaded. Extensions are plug-ins into doca-bench; they do not replace it, they do not have a standalone CLI, they do not measure anything without the parent invoking them. Every question on this skill presupposes the parent.
  • doca-version — the DOCA_EXPERIMENTAL surface is versioned with DOCA; the extension's soversion matches the DOCA release per the shipped meson.build. The four-way version match applies; rebuilding the extension across DOCA upgrades is the rule, not the exception.
  • doca-gpunetio — when the extension is GPU-side and uses GPUNetIO RX / TX queues like the reference doca_bench_cuda. Route the GPUNetIO semantics there.
  • doca-setup — DOCA install posture (does doca-bench exist? does the doca_bench_cuda_impl reference library exist? is the CUDA toolchain installed when needed?).
  • doca-debug — the cross-cutting debug ladder for env-side issues (dynamic linker, library search path, CUDA driver / toolkit, firmware).
  • doca-public-knowledge-map — routing to the public DOCA Bench / DOCA GPUNetIO pages on docs.nvidia.com and the release notes for the documented extension lifecycle / discovery mechanism.
  • doca-structured-tools-contract — the agent's detect → prefer → fall back → report contract for the structured helpers (doca-env --json, doca-capability-snapshot, version-matrix.json) the build / load preconditions rely on.
  • doca-hardware-safety — the canonical hardware-safety meta-policy that CAPABILITIES.md ## Safety policy overlays. Extensions are external code loaded into doca-bench; the safety implications of loading experimental code into a benchmark that touches the dataplane / device are real.

This skill assumes the user has built shared libraries on Linux before and knows what a Meson build is. Background material on those topics belongs in the toolchain docs, not in this skill.

nvidia의 다른 스킬

compileiq-debug
nvidia
무언가 잘못되었을 때 사용: Search()가 멈추거나, 모든 평가가 INVALID_SCORE를 반환하거나, 점수가 개선되지 않거나, 모든 설정이 동일한 숫자를 반환하거나, ptxas 오류 등이 발생할 때
create-github-pr
nvidia
gh CLI를 사용하여 GitHub 풀 리퀘스트를 생성합니다. 사용자가 새 PR을 만들거나, 코드 리뷰를 제출하거나, 풀 리퀘스트를 열고자 할 때 사용합니다. 트리거 키워드 -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
다른 열린 이슈들을 스캔하여 주어진 PR이 함께 수정하거나 실수로 망가뜨릴 수 있는 이슈를 찾습니다. 인접 수정 기회와 모순 위험을 file:line…과 함께 출력합니다.
fhir-basics
nvidia
에이전트에게 FHIR R4 API의 작동 방식, 사용 가능한 리소스, 검색 매개변수를 사용한 쿼리 방법, 모든 응답 형식을 올바르게 파싱하는 방법을 가르칩니다…
compileiq-validate-result
nvidia
검색이 완료된 후, 속도 향상을 청구하거나 ACF를 발송하기 전에 사용합니다. dump_results CSV를 로드하고, 상위 K개 후보(단일 목표)를 추출합니다…
changelog-audit
nvidia
릴리스 전에 Warp CHANGELOG.md를 감사합니다: 누락된 항목 복구, 사용자 영향별 정렬, 항목 언어 다듬기, 줄 바꿈, (릴리스 브랜치 모드) 비교 업데이트…
maintain-dynamic-plugins
nvidia
NeMo Relay 동적 플러그인 로더, 매니페스트, Rust 네이티브 SDK, gRPC 워커 프로토콜, Python 워커 SDK, 문서, 테스트 및 릴리스 워크플로 커버리지를 유지 관리합니다.
dgx-diagnose
nvidia
일반적인 DGX Station GB300 문제 진단 — CUDA 충돌, 잘못된 GPU 타겟팅, vLLM/SGLang 컨테이너 버그, MIG 상태 문제, NVLink/Fabric Manager 오류,…