doca-compress

작성자: nvidia

BlueField DPU, ConnectX NIC 또는 DOCA가 있는 호스트에서 실습형 DOCA Compress 프로그래밍을 위해 이 스킬을 사용하여 compress-deflate, decompress-deflate 등을 활성화합니다.

npx skills add https://github.com/nvidia/skills --skill doca-compress

DOCA Compress

Where to start: This skill assumes DOCA is already installed and the user is doing hands-on Compress work (bulk DEFLATE compression or decompression) on a BlueField / ConnectX / host with DOCA. Open TASKS.md if the user wants to do something (configure / build / modify / run / test / debug); open CAPABILITIES.md when the question is what can DOCA Compress express on this version. If the user has not installed DOCA yet, route to doca-setup first. If the user is asking "should I even offload this compression to the accelerator?", the size-threshold path-selection rule in CAPABILITIES.md ## Capabilities and modes is the first stop — bulk compress is the canonical fit, tiny one-shot is not.

Example questions this skill answers well

The CLASSES of DOCA Compress questions this skill is built to answer, each with one worked example. The agent should treat the class as the load-bearing piece — the worked example is a single instance.

  • "Should I offload this compression to DOCA Compress, or just zlib/zstd on the CPU?" — worked example: "I am compressing a 4 MiB log buffer before writing it to storage; is doca-compress worth the setup vs zlib on the CPU?". Answered by the size-threshold path-selection rule in CAPABILITIES.md ## Capabilities and modes
  • "Does my device support the compress (or decompress) task I want, and how big a buffer can it move per submission?" — worked example: "is doca_compress_task_compress_deflate on this BlueField, and what is the max source size per task?". Answered by the per-task capability-query rule (doca_compress_cap_task_compress_deflate_is_supported, _decompress_deflate_is_supported, the matching _get_max_buf_size queries) in CAPABILITIES.md ## Capabilities and modes
  • "I just need to decompress incoming network data — is that a valid standalone use of doca-compress?" — worked example: "my client receives DEFLATE-compressed payloads and never produces any compressed output of its own". Answered by the "decompression-only" note in CAPABILITIES.md ## Capabilities and modes task-type table + the per-task configuration matrix in TASKS.md ## configure step 5.
  • "What permissions does the source / destination mmap need?" — worked example: "my doca_compress_task_compress_deflate returns DOCA_ERROR_NOT_PERMITTED". Answered by the permission matrix in CAPABILITIES.md ## Safety policy
  • "Is this DOCA Compress API available on my installed DOCA version?" — worked example: "is doca_compress_task_decompress_deflate in the DOCA I have installed?". Answered by the version-compatibility overlay in CAPABILITIES.md ## Version compatibility, which cross-links the canonical detection chain in doca-version and adds the Compress-specific "discover per-task support, do not assume" bullets.
  • "What does this DOCA_ERROR_* from a Compress call mean and which layer caused it?" — worked example: "DOCA_ERROR_INVALID_VALUE on doca_compress_task_compress_deflate_alloc_init". Answered by the Compress overlay on the cross-library taxonomy in CAPABILITIES.md ## Error taxonomy

Audience

This skill serves external developers building applications that consume the DOCA Compress library — i.e., users whose code calls doca_compress_* (directly in C/C++, or through FFI/bindings from another language) to offload bulk DEFLATE compression or decompression onto a BlueField DPU or ConnectX accelerator. It is not for NVIDIA developers contributing to DOCA Compress itself.

Language scope. DOCA Compress ships as a C library with pkg-config module name doca-compress. The shipped samples are written in C. C and C++ consumers are the canonical case and the worked examples in TASKS.md assume that path. Other-language consumers (Rust, Go, Python, …) consume the same *.so through FFI or language-specific bindings; the skill's contribution in that case is to keep the lifecycle, capability-discovery, permission, error-taxonomy, and compress-vs-decompress guidance language-neutral, and to route the agent to the public C ABI as the authoritative surface that any wrapper will eventually call.

When to load this skill

Load this skill when the user is doing hands-on DOCA Compress work, in any language. Concretely:

  • Initializing a doca_compress context on a doca_dev and configuring at least one task type (doca_compress_task_compress_deflate and/or doca_compress_task_decompress_deflate) before doca_ctx_start().
  • Choosing between the compress task and the decompress task for the user's data flow. The two task types are independent: a consumer that only decompresses inbound DEFLATE payloads is a valid shape and does not need to enable the compress task.
  • Setting permissions on doca_mmap correctly for the source buffer (DOCA_ACCESS_FLAG_LOCAL_READ_ONLY at minimum) and the destination buffer (DOCA_ACCESS_FLAG_LOCAL_READ_WRITE).
  • Sizing the source buffer against the per-task doca_compress_cap_task_*_get_max_buf_size(devinfo) ceiling and the destination buffer against the worst-case output size the algorithm can produce on this input.
  • Checking which task types this device's accelerator advertises via doca_compress_cap_task_compress_deflate_is_supported and doca_compress_cap_task_decompress_deflate_is_supported against the active doca_devinfo.
  • Validating against a round-trip (compress → decompress, or decompress → compare against a known DEFLATE-encoded fixture) before pushing bulk input through the accelerator.
  • Deciding whether to offload at all — the size-threshold rule in CAPABILITIES.md ## Capabilities and modes says doca-compress is the right answer only for bulk inputs (rule of thumb: ≥ a few KiB); below that, CPU compression beats the DMA-to-accelerator round-trip.
  • Debugging a DOCA_ERROR_* returned from a Compress call (lifecycle vs. buffer-sizing vs. permission vs. unsupported-task) and the task-completion event on the progress engine.
  • Designing or extending non-C bindings (Rust, Go, Python, …) that wrap the Compress C ABI — for the lifecycle, permission, capability, and compress-vs-decompress rules the wrapper must honor.

Do not load this skill for general DOCA orientation, install of DOCA itself, non-DEFLATE compression libraries on CPU (use zlib / zstd / similar), or other DOCA libraries. For those, use doca-public-knowledge-map.

What this skill provides

This is a thin loader. The body keeps only the orientation needed to pick the right next file. The substantive Compress-specific material lives in two companion files:

  • CAPABILITIES.md — what DOCA Compress can express on this version: the four task types (compress-deflate, decompress-deflate, decompress-lz4-stream, and decompress-lz4-block, each independently capability-gated), the per-task capability-query surface (doca_compress_cap_* for task support and per-submission max buffer size), the Compress error taxonomy (mapped onto the cross-library DOCA_ERROR_* set), the observability surface (per-task completion events on the progress engine), the safety policy that gates source / destination mmap permission decisions, and the size-threshold path-selection rule (when to use doca-compress versus CPU compression or doca-dma for the no-compression copy case).
  • TASKS.md — step-by-step workflows for the six in-scope Compress verbs: configure, build, modify, run, test, debug. Plus a ## rollback overlay (Compress-specific five-step teardown that drains in-flight tasks before doca_ctx_stop, unregisters mmap regions in reverse-register order, and re-verifies the device path with the round-trip smoke) and the 5-phase universal debug-loop instantiation appended to ## debug. Plus a Deferred task verbs block that points out-of-scope questions at the right next skill.

The skill assumes a host or BlueField where DOCA is already installed at the standard location and the user has the privileges their public install profile expects. It does not cover installing DOCA — that path goes through doca-setup.

What this skill deliberately does not ship

This skill is agent guidance, not a samples or templates bundle. To keep the boundary clean, it deliberately does not contain — and pull requests should not add:

  • Pre-written DOCA Compress application source code, in any language. The verified Compress source code is the shipped C samples at /opt/mellanox/doca/samples/doca_compress/, plus the File Compression reference application linked from the public DOCA Compress guide. The agent's job is to route the user to those files and prescribe a minimum-diff modification on them via the universal modify-a-sample workflow in doca-programming-guide, layered with the Compress-specific overrides in TASKS.md ## modify.
  • Pre-encoded DEFLATE fixtures for arbitrary inputs. The skill tells the agent to use a round-trip validation (compress → decompress and compare to the original, or decompress a known DEFLATE blob produced by a published CPU encoder like zlib) as the known-good smoke; it does not ship a fixture bank of its own.
  • Standalone build manifests (meson.build, CMakeLists.txt, Cargo.toml, …) parked inside the skill. The agent constructs the build manifest in the user's project directory against the user's installed DOCA, where pkg-config --modversion doca-compress is the source of truth.
  • A samples/, bindings/, or reference/ subtree of any kind. A mock or incomplete artifact in this skill's tree, even one labeled "reference", is misleading: users will read it as buildable.

Loading order

  1. Read this SKILL.md first to confirm the user's question is in scope.
  2. For the Compress capability matrix, the compress / decompress task split (including decompression-only as a standalone shape), per-task capability-query rules, permission matrix, error taxonomy, observability, and safety / path-selection policy, see CAPABILITIES.md.
  3. For step-by-step workflows — configure, build, modify, run, test, debug — see TASKS.md.

Both companion files cross-link to each other, doca-version for the canonical version-handling rules, and doca-public-knowledge-map whenever the right answer is "look it up in the public docs or the installed package layout" rather than "Compress-specific guidance".

Related skills

  • doca-public-knowledge-map — the routing table for every public DOCA documentation source and the on-disk layout of an installed DOCA package. The DOCA Compress page lives at docs.nvidia.com/doca/sdk/DOCA-Compress/; the File Compression reference application is the canonical worked example.
  • doca-setup — env preparation, install verification, and the I have no install yet path with the public NGC DOCA container. This skill assumes its preconditions are satisfied.
  • doca-version — canonical DOCA version-handling rules. This skill's ## Version compatibility cross-links the four-way match rule and adds only the Compress-specific "discover per-task support + per-task max buffer size via cap query" overlay.
  • doca-structured-tools-contract — the bundle's structured-tools precedence rule (detect / prefer / fall back / report). The Command appendix in TASKS.md honors this contract.
  • doca-programming-guide — general DOCA programming patterns shared by every library: the canonical pkg-config + meson build pattern, the universal modify-a-shipped-sample first-app workflow, the universal lifecycle, the cross-library DOCA_ERROR_* taxonomy, and the program-side debug order. This skill layers Compress specifics on top.
  • doca-dma — the right library when the data flow is a pure mmap-to-mmap copy with no compression required. This skill's path-selection rule routes to DMA when Compress is not the answer (e.g. the user only wants to move bytes, not encode them).
  • doca-debug — the cross-cutting debug ladder (install / version / build / link / runtime / program / driver). Compress-specific debug (task-not-supported, source-buffer-too-large, destination-buffer-too-small) overlays on top of that ladder.

nvidia의 다른 스킬

compileiq-debug
nvidia
무언가 잘못되었을 때 사용: Search()가 멈추거나, 모든 평가가 INVALID_SCORE를 반환하거나, 점수가 개선되지 않거나, 모든 설정이 동일한 숫자를 반환하거나, ptxas 오류 등이 발생할 때
create-github-pr
nvidia
gh CLI를 사용하여 GitHub 풀 리퀘스트를 생성합니다. 사용자가 새 PR을 만들거나, 코드 리뷰를 제출하거나, 풀 리퀘스트를 열고자 할 때 사용합니다. 트리거 키워드 -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
다른 열린 이슈들을 스캔하여 주어진 PR이 함께 수정하거나 실수로 망가뜨릴 수 있는 이슈를 찾습니다. 인접 수정 기회와 모순 위험을 file:line…과 함께 출력합니다.
fhir-basics
nvidia
에이전트에게 FHIR R4 API의 작동 방식, 사용 가능한 리소스, 검색 매개변수를 사용한 쿼리 방법, 모든 응답 형식을 올바르게 파싱하는 방법을 가르칩니다…
compileiq-validate-result
nvidia
검색이 완료된 후, 속도 향상을 청구하거나 ACF를 발송하기 전에 사용합니다. dump_results CSV를 로드하고, 상위 K개 후보(단일 목표)를 추출합니다…
changelog-audit
nvidia
릴리스 전에 Warp CHANGELOG.md를 감사합니다: 누락된 항목 복구, 사용자 영향별 정렬, 항목 언어 다듬기, 줄 바꿈, (릴리스 브랜치 모드) 비교 업데이트…
maintain-dynamic-plugins
nvidia
NeMo Relay 동적 플러그인 로더, 매니페스트, Rust 네이티브 SDK, gRPC 워커 프로토콜, Python 워커 SDK, 문서, 테스트 및 릴리스 워크플로 커버리지를 유지 관리합니다.
dgx-diagnose
nvidia
일반적인 DGX Station GB300 문제 진단 — CUDA 충돌, 잘못된 GPU 타겟팅, vLLM/SGLang 컨테이너 버그, MIG 상태 문제, NVLink/Fabric Manager 오류,…