cicd

par nvidia

Référence CI/CD pour NeMo-RL. Couvre la structure du pipeline GitHub Actions, le déclenchement du CI via /ok to test, et l'investigation des échecs du CI.

npx skills add https://github.com/NVIDIA-NeMo/RL --skill cicd

CI/CD Guide


How CI Works

NeMo-RL CI runs on GitHub Actions. Workflows live in .github/workflows/.

The main test workflow triggers on pushes to pull-request/<number> branches. These branches are created automatically by copy-pr-bot when a contributor pushes to their fork and the PR receives a trust signal:

  • The commit is GPG-signed by a maintainer, or
  • A maintainer posts /ok to test <full-sha> as a PR comment.

Triggering CI

After pushing a new commit to your PR, trigger CI with:

/ok to test <full-commit-sha>

Use git rev-parse HEAD (not the short form) to get the full SHA.

Re-triggering after new commits: each /ok to test <sha> is specific to that SHA. After pushing additional commits, post a new comment with the updated SHA.


CI Labels

Every PR must have exactly one CI:* label — the quality-check job stays red until one is attached. Labels control which test tier runs and whether a new container image is built.

LabelWhat runsContainer
CI:docsDoc tests onlyReuses main container
CI:LfastFast test subsetReuses main container
CI:L0Unit tests + docs + lintBuilds new image
CI:L1L0 + functional testsBuilds new image
CI:L2L1 + convergence testsBuilds new image
Skip CICDNothing (skips all tests)

Default on merge group / push to main: L1.

Which label to attach when opening a PR:

Changed paths / nature of changeLabel
Docs only (docs/, *.md, docstrings)CI:docs
Trivial fix, no logic changeCI:Lfast
New code, bug fix, refactorCI:L0
Changes that could affect model behaviourCI:L1
Changes that could affect convergenceCI:L2

CI Failure Investigation

# List recent workflow runs for the PR
gh run list --repo NVIDIA-NeMo/RL --branch "pull-request/<pr-number>"

# View failing run summary
gh run view <run-id> --repo NVIDIA-NeMo/RL

# Stream failing job output
gh run view <run-id> --repo NVIDIA-NeMo/RL --log-failed

Common failure patterns:

SymptomLikely causeFix
CI never startsNo trust signal or copy-pr-bot not triggeredPost /ok to test <sha>
semantic-pull-request failsPR title doesn't follow Conventional CommitsFix PR title; see contributing skill
Linting failsStyle violationRun uv run ruff check --fix . && uv run ruff format .
Unit test failureCode regression or missing dependencyReproduce locally; see testing skill

Plus de skills de nvidia

compileiq-debug
nvidia
Utilisez quand quelque chose ne va pas : Search() bloque, toutes les évaluations retournent INVALID_SCORE, les scores ne s'améliorent pas, chaque configuration retourne le même nombre, erreurs ptxas…
create-github-pr
nvidia
Créer des pull requests GitHub en utilisant l'interface en ligne de commande gh. Utiliser lorsque l'utilisateur souhaite créer une nouvelle PR, soumettre du code pour révision, ou ouvrir une pull request. Mots-clés de déclenchement -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Analyse les autres problèmes ouverts pour trouver ceux qu’une PR donnée pourrait également corriger ou casser accidentellement. Génère des opportunités de correctifs adjacents et des risques de contradiction avec fichier:ligne…
fhir-basics
nvidia
Apprend aux agents comment fonctionnent les API FHIR R4, quelles ressources sont disponibles, comment les interroger avec des paramètres de recherche, et comment analyser correctement tous les formats de réponse…
compileiq-validate-result
nvidia
Utiliser APRÈS qu'une recherche soit terminée et AVANT de réclamer un accélérateur ou d'expédier un ACF. Charge le CSV dump_results, extrait les K meilleurs candidats (mono-objectif)…
changelog-audit
nvidia
Auditer le CHANGELOG.md de Warp avant une publication : récupérer les entrées perdues, trier par impact utilisateur, affiner le langage des entrées, ajuster les retours à la ligne et (en mode branche de publication) mettre à jour la comparaison…
maintain-dynamic-plugins
nvidia
Maintenir les chargeurs de plugins dynamiques NeMo Relay, les manifestes, les SDK natifs Rust, le protocole worker gRPC, le SDK worker Python, la documentation, les tests et la couverture du workflow de publication
dgx-diagnose
nvidia
Diagnostiquer les problèmes courants du DGX Station GB300 — plantages CUDA, ciblage incorrect du GPU, bugs de conteneur vLLM/SGLang, problèmes d'état MIG, erreurs NVLink/Fabric Manager,…