nvrx-attr

โดย nvidia

ชั้นการจัดเรียงเหนือโมดูลการระบุแหล่งที่มาของ nvidia_resiliency_ext ให้การวิเคราะห์บันทึก การวิเคราะห์ fr และข้อเสนอแนะการฉีดข้อบกพร่องที่เน้น Megatron-LM…

npx skills add https://github.com/nvidia/nvidia-resiliency-ext --skill nvrx-attr

Attribution Skills

High-level orchestration layer over the nvidia_resiliency_ext.attribution modules. Each subdirectory is a self-contained skill with its own SKILL.md and helper scripts.

Skills

DirectoryPurposeEntry point
log-analysis/Analyze SLURM job logs for failure root-cause and restart decisionsNVRxLogAnalyzer (nvrx_logsage.py)
fr-analysis/Analyze NCCL flight-recorder dumps for collective-hang root-causeCollectiveAnalyzer (fr_attribution.py)
fault-injection-loop/Run a batched SLURM fault-injection feedback loop and score attribution accuracyprepare_node_alloc.sh / watch_and_analyze.sh

How skills relate to the library

src/nvidia_resiliency_ext/
├── attribution/
│   ├── log_analyzer/nvrx_logsage.py      ← log-analysis implementation
│   ├── trace_analyzer/fr_attribution.py  ← fr-analysis implementation
│   ├── analyzer/engine.py                ← combined orchestration entry point
│   └── combined_log_fr/                  ← optional log + FR fusion
└── skills/
    └── nvrx-attr/                        ← this skill bundle
        ├── log-analysis/
        ├── fr-analysis/
        └── fault-injection-loop/

The Analyzer (analyzer/engine.py) is the recommended entry point when you need request coalescing, result caching, or the combined LOG_AND_TRACE pipeline. Use the individual skills when you want to run one analysis type directly without the full coalescing stack.

Common prerequisites

  • LLM_API_KEY environment variable, LLM_API_KEY_FILE, or ~/.llm_api_key
  • Package installed: pip install 'nvidia-resiliency-ext[attribution]' or pip install -e '.[attribution]' from repo root
  • The fault-injection loop has only been validated with Megatron-LM training scripts

Fault-Loop Local Setup

Before using fault-injection-loop/, create the local config file from the tracked template and fill in your site-specific values:

cp scripts/user.env.example scripts/user.env

The feedback-loop scripts require src/nvidia_resiliency_ext/skills/nvrx-attr/scripts/user.env to exist at runtime. Keep user.env local and untracked.

Skills เพิ่มเติมจาก nvidia

compileiq-debug
nvidia
ใช้เมื่อมีบางอย่างผิดปกติ: Search() ค้าง, การประเมินทั้งหมดคืนค่า INVALID_SCORE, คะแนนไม่ดีขึ้น, ทุกคอนฟิกคืนค่าเลขเดียวกัน, ข้อผิดพลาด ptxas…
create-github-pr
nvidia
สร้างคำขอดึงข้อมูล GitHub โดยใช้ gh CLI ใช้เมื่อผู้ใช้ต้องการสร้าง PR ใหม่ ส่งโค้ดเพื่อตรวจสอบ หรือเปิดคำขอดึงข้อมูล คำหลักที่ใช้เรียก -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
สแกน issue อื่นๆ ที่เปิดอยู่เพื่อค้นหาว่า PR ที่กำหนดอาจแก้ไขหรือทำให้เสียโดยไม่ได้ตั้งใจ แสดงผลโอกาสในการแก้ไขที่เกี่ยวข้องและความเสี่ยงที่ขัดแย้งกันพร้อมไฟล์:บรรทัด…
fhir-basics
nvidia
สอนให้เอเจนต์เข้าใจการทำงานของ FHIR R4 API ทรัพยากรที่มีให้ วิธีค้นหาด้วยพารามิเตอร์ค้นหา และวิธีแยกวิเคราะห์รูปแบบการตอบกลับทั้งหมดอย่างถูกต้อง…
compileiq-validate-result
nvidia
ใช้หลังจากที่การค้นหาเสร็จสิ้น และก่อนที่จะอ้างสิทธิ์การเร่งความเร็วหรือจัดส่ง ACF โหลดไฟล์ CSV dump_results แยกผู้สมัคร K อันดับแรก (วัตถุประสงค์เดียว)…
changelog-audit
nvidia
ตรวจสอบ Warp CHANGELOG.md ก่อนปล่อย: กู้คืนรายการที่สูญหาย จัดเรียงตามผลกระทบต่อผู้ใช้ ปรับปรุงภาษาในรายการ จัดบรรทัด และ (ในโหมดสาขาปล่อย) เปรียบเทียบการเพิ่มเวอร์ชัน…
maintain-dynamic-plugins
nvidia
ดูแล NeMo Relay dynamic plugin loaders, manifests, Rust native SDKs, gRPC worker protocol, Python worker SDK, เอกสาร, การทดสอบ และความครอบคลุมของเวิร์กโฟลว์การเผยแพร่
dgx-diagnose
nvidia
วินิจฉัยปัญหาทั่วไปของ DGX Station GB300 — CUDA ล่ม, การกำหนดเป้าหมาย GPU ผิด, บั๊กคอนเทนเนอร์ vLLM/SGLang, ปัญหาสถานะ MIG, ข้อผิดพลาด NVLink/Fabric Manager,…