cutedsl-kernel-integration

von nvidia

Verwenden Sie bei der Integration eines CuTeDSL/CUTE DSL-Kernels in das cuDNN Frontend als reines Python-Frontend-API, einschließlich APIBase-Wrapper, lazy cudnn-Exporte, optional…

npx skills add https://github.com/nvidia/cudnn-frontend --skill cutedsl-kernel-integration

CuTeDSL Kernel Integration

Use this skill to add or update a CuTeDSL frontend-only API in cuDNN Frontend. The goal is a complete integration: Python API, wrapper, exports, docs, and tests.

Before Editing

  1. Inspect the current repo state and avoid overwriting unrelated changes.
  2. Confirm every original source file needed for the integration is available. If a source file is missing, report that gap instead of inferring its contract from a related kernel.
  3. Record source provenance when it is available: upstream URL, local source path, commit, and which files map to public API modules versus private helpers.
  4. Classify the kernel before choosing a template:
    • Kernel family: dense GEMM, GEMM fusion, grouped GEMM, discrete grouped GEMM, MoE, attention, sparse attention, or another frontend-only API family.
    • Execution topology: single kernel, paired forward/backward APIs, multi-kernel orchestrator, helper-kernel setup, distributed/runtime-coordinated execution, or internal scheduler.
    • Public surface: class API, high-level wrapper, returned tensors, optional outputs, workspace ownership, and import/export namespace.
    • Internal support: source helper modules, schedulers, metadata utilities, and generated descriptors that must stay private to the package.
    • Architecture variant: whether the public API needs transparent dispatch to an alternate CuTeDSL module for a newer GPU (for example Rubin sm107 vs the default SM100 kernel). Keep the public class and wrapper unchanged when dispatch is internal.
  5. Read references/integration-pattern.md for the detailed repo conventions before implementing.

Integration Workflow

  1. Add or update the operation package under the closest existing family, such as python/cudnn/<operation>/, python/cudnn/gemm/cutedsl/dense/<operation>/, python/cudnn/gemm/cutedsl/grouped/<operation>/, python/cudnn/gemm/cutedsl/discrete_grouped/<operation>/, or python/cudnn/sdpa/<direction>/.
  2. Implement the class API by extending APIBase; keep constructor descriptors, check_support(), compile(), and execute() consistent with the closest template.
  3. Add a high-level wrapper that allocates outputs, caches/reuses compiled kernels where the template does, and returns a TupleDict.
  4. Export the public class and wrapper through the operation/family __init__.py files and _LAZY_OPTIONAL_IMPORTS in python/cudnn/__init__.py.
  5. Reuse the existing CuTeDSL dependencies in [project] dependencies unless the new kernel truly needs an additional package. The cutedsl extra now holds only cuda-python.
  6. Add FE OSS documentation and update the relevant overview or operation index links.
  7. Add tests under test/python/<operation>/cutedsl/, including support validation and numerical/reference coverage when executable.
  8. For grouped/discrete/MoE/SDPA kernels, preserve the source helper and scheduler topology; shared helper modules should be internal package files, not public cudnn exports.
  9. When an existing SM100 kernel needs a Rubin (sm107) variant, follow the architecture-dispatch pattern in references/integration-pattern.md instead of exposing a new public API. Current examples: grouped_gemm_quant, grouped_gemm_glu, and grouped_gemm_dglu.

Verification

  • Run focused formatting or tests for the files changed.
  • At minimum for skill-only edits, verify this SKILL.md has valid frontmatter and all referenced paths exist.
  • For kernel integrations, run the relevant pytest test/python/<operation>/cutedsl/test_<operation>.py target when the environment has the required GPU and optional dependencies; otherwise report the skipped verification explicitly.
  • For architecture-dispatch work, also run pytest test/python/gemm/cutedsl/test_rubin_kernel_dispatch.py. On Rubin hardware, the existing FE API e2e tests for the affected operation should still pass without API changes.

Mehr Skills von nvidia

fhir-basics
nvidia
Bringt Agenten bei, wie FHIR R4 APIs funktionieren, welche Ressourcen verfügbar sind, wie man sie mit Suchparametern abfragt und wie man alle Antwortformate korrekt parst…
compileiq-validate-result
nvidia
Verwende NACH Abschluss einer Suche und VOR dem Einfordern eines Speedups oder dem Versand eines ACF. Lädt die dump_results CSV, extrahiert Top-K-Kandidaten (Einzelziel)...
changelog-audit
nvidia
Auditiere die CHANGELOG.md vor einem Release: stelle verlorene Einträge wieder her, sortiere nach Benutzerauswirkung, verfeinere die Sprache der Einträge, führe Zeilenumbrüche durch und (im Release-Branch-Modus) erhöhe die Vergleichsnummer…
dgx-diagnose
nvidia
Diagnostizieren Sie häufige DGX Station GB300-Probleme – CUDA-Abstürze, falsche GPU-Zuweisung, vLLM/SGLang-Container-Fehler, MIG-Status-Probleme, NVLink/Fabric-Manager-Fehler,…
aicr-managing-openvex
nvidia
Use when adding, updating, or removing CVE/GHSA suppressions in `.openvex.json` — the OpenVEX document consumed by the daily image vulnerability scan workflow.…
aicr-creating-slide-decks
nvidia
Verwenden Sie beim Erstellen eines eigenständigen HTML-Foliensatzes oder visuellen Gesprächspunkts für ein technisches Konzept oder einen Workflow (z. B. eine Demos/*.html-Datei) — angezeigt im Vollbildmodus oder…
aicr-creating-guided-demos
nvidia
Erstellt ein interaktives geführtes Demo-Skript (demos/*.sh), live oder im eigenen Tempo, mit dem Muster Frame → Tell → Show → Close. Wird bei „Demo-Skript", „geführt…" ausgelöst.
aicr-analyzing-snapshots
nvidia
Verwenden Sie bei der Analyse einer AICR-Snapshot-YAML-Datei, bei der Überprüfung des Cluster-Zustands, beim Vergleich von Provider-Eigenschaften, beim Extrahieren von GPU-/Netzwerktopologie-Erkenntnissen oder…