cicd

โดย nvidia

ข้อมูลอ้างอิง CI/CD สำหรับ NeMo-RL ครอบคลุมโครงสร้างไปป์ไลน์ GitHub Actions การทริกเกอร์ CI ผ่าน /ok to test และการตรวจสอบความล้มเหลวของ CI

npx skills add https://github.com/NVIDIA-NeMo/RL --skill cicd

CI/CD Guide


How CI Works

NeMo-RL CI runs on GitHub Actions. Workflows live in .github/workflows/.

The main test workflow triggers on pushes to pull-request/<number> branches. These branches are created automatically by copy-pr-bot when a contributor pushes to their fork and the PR receives a trust signal:

  • The commit is GPG-signed by a maintainer, or
  • A maintainer posts /ok to test <full-sha> as a PR comment.

Triggering CI

After pushing a new commit to your PR, trigger CI with:

/ok to test <full-commit-sha>

Use git rev-parse HEAD (not the short form) to get the full SHA.

Re-triggering after new commits: each /ok to test <sha> is specific to that SHA. After pushing additional commits, post a new comment with the updated SHA.


CI Labels

Every PR must have exactly one CI:* label — the quality-check job stays red until one is attached. Labels control which test tier runs and whether a new container image is built.

LabelWhat runsContainer
CI:docsDoc tests onlyReuses main container
CI:LfastFast test subsetReuses main container
CI:L0Unit tests + docs + lintBuilds new image
CI:L1L0 + functional testsBuilds new image
CI:L2L1 + convergence testsBuilds new image
Skip CICDNothing (skips all tests)

Default on merge group / push to main: L1.

Which label to attach when opening a PR:

Changed paths / nature of changeLabel
Docs only (docs/, *.md, docstrings)CI:docs
Trivial fix, no logic changeCI:Lfast
New code, bug fix, refactorCI:L0
Changes that could affect model behaviourCI:L1
Changes that could affect convergenceCI:L2

CI Failure Investigation

# List recent workflow runs for the PR
gh run list --repo NVIDIA-NeMo/RL --branch "pull-request/<pr-number>"

# View failing run summary
gh run view <run-id> --repo NVIDIA-NeMo/RL

# Stream failing job output
gh run view <run-id> --repo NVIDIA-NeMo/RL --log-failed

Common failure patterns:

SymptomLikely causeFix
CI never startsNo trust signal or copy-pr-bot not triggeredPost /ok to test <sha>
semantic-pull-request failsPR title doesn't follow Conventional CommitsFix PR title; see contributing skill
Linting failsStyle violationRun uv run ruff check --fix . && uv run ruff format .
Unit test failureCode regression or missing dependencyReproduce locally; see testing skill

Skills เพิ่มเติมจาก nvidia

compileiq-debug
nvidia
ใช้เมื่อมีบางอย่างผิดปกติ: Search() ค้าง, การประเมินทั้งหมดคืนค่า INVALID_SCORE, คะแนนไม่ดีขึ้น, ทุกคอนฟิกคืนค่าเลขเดียวกัน, ข้อผิดพลาด ptxas…
create-github-pr
nvidia
สร้างคำขอดึงข้อมูล GitHub โดยใช้ gh CLI ใช้เมื่อผู้ใช้ต้องการสร้าง PR ใหม่ ส่งโค้ดเพื่อตรวจสอบ หรือเปิดคำขอดึงข้อมูล คำหลักที่ใช้เรียก -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
สแกน issue อื่นๆ ที่เปิดอยู่เพื่อค้นหาว่า PR ที่กำหนดอาจแก้ไขหรือทำให้เสียโดยไม่ได้ตั้งใจ แสดงผลโอกาสในการแก้ไขที่เกี่ยวข้องและความเสี่ยงที่ขัดแย้งกันพร้อมไฟล์:บรรทัด…
fhir-basics
nvidia
สอนให้เอเจนต์เข้าใจการทำงานของ FHIR R4 API ทรัพยากรที่มีให้ วิธีค้นหาด้วยพารามิเตอร์ค้นหา และวิธีแยกวิเคราะห์รูปแบบการตอบกลับทั้งหมดอย่างถูกต้อง…
compileiq-validate-result
nvidia
ใช้หลังจากที่การค้นหาเสร็จสิ้น และก่อนที่จะอ้างสิทธิ์การเร่งความเร็วหรือจัดส่ง ACF โหลดไฟล์ CSV dump_results แยกผู้สมัคร K อันดับแรก (วัตถุประสงค์เดียว)…
changelog-audit
nvidia
ตรวจสอบ Warp CHANGELOG.md ก่อนปล่อย: กู้คืนรายการที่สูญหาย จัดเรียงตามผลกระทบต่อผู้ใช้ ปรับปรุงภาษาในรายการ จัดบรรทัด และ (ในโหมดสาขาปล่อย) เปรียบเทียบการเพิ่มเวอร์ชัน…
maintain-dynamic-plugins
nvidia
ดูแล NeMo Relay dynamic plugin loaders, manifests, Rust native SDKs, gRPC worker protocol, Python worker SDK, เอกสาร, การทดสอบ และความครอบคลุมของเวิร์กโฟลว์การเผยแพร่
dgx-diagnose
nvidia
วินิจฉัยปัญหาทั่วไปของ DGX Station GB300 — CUDA ล่ม, การกำหนดเป้าหมาย GPU ผิด, บั๊กคอนเทนเนอร์ vLLM/SGLang, ปัญหาสถานะ MIG, ข้อผิดพลาด NVLink/Fabric Manager,…