cicd

por nvidia

Referência de CI/CD para NeMo-RL. Abrange a estrutura do pipeline do GitHub Actions, acionamento de CI via /ok to test e investigação de falhas de CI.

npx skills add https://github.com/nvidia-nemo/rl --skill cicd

CI/CD Guide


How CI Works

NeMo-RL CI runs on GitHub Actions. Workflows live in .github/workflows/.

The main test workflow triggers on pushes to pull-request/<number> branches. These branches are created automatically by copy-pr-bot when a contributor pushes to their fork and the PR receives a trust signal:

  • The commit is GPG-signed by a maintainer, or
  • A maintainer posts /ok to test <full-sha> as a PR comment.

Triggering CI

After pushing a new commit to your PR, trigger CI with:

/ok to test <full-commit-sha>

Use git rev-parse HEAD (not the short form) to get the full SHA.

Re-triggering after new commits: each /ok to test <sha> is specific to that SHA. After pushing additional commits, post a new comment with the updated SHA.


CI Labels

Every PR must have exactly one CI:* label — the quality-check job stays red until one is attached. Labels control which test tier runs and whether a new container image is built.

LabelWhat runsContainer
CI:docsDoc tests onlyReuses main container
CI:LfastFast test subsetReuses main container
CI:L0Unit tests + docs + lintBuilds new image
CI:L1L0 + functional testsBuilds new image
CI:L2L1 + convergence testsBuilds new image
Skip CICDNothing (skips all tests)—

Default on merge group / push to main: L1.

Which label to attach when opening a PR:

Changed paths / nature of changeLabel
Docs only (docs/, *.md, docstrings)CI:docs
Trivial fix, no logic changeCI:Lfast
New code, bug fix, refactorCI:L0
Changes that could affect model behaviourCI:L1
Changes that could affect convergenceCI:L2

CI Failure Investigation

# List recent workflow runs for the PR
gh run list --repo NVIDIA-NeMo/RL --branch "pull-request/<pr-number>"

# View failing run summary
gh run view <run-id> --repo NVIDIA-NeMo/RL

# Stream failing job output
gh run view <run-id> --repo NVIDIA-NeMo/RL --log-failed

Common failure patterns:

SymptomLikely causeFix
CI never startsNo trust signal or copy-pr-bot not triggeredPost /ok to test <sha>
semantic-pull-request failsPR title doesn't follow Conventional CommitsFix PR title; see contributing skill
Linting failsStyle violationRun uv run ruff check --fix . && uv run ruff format .
Unit test failureCode regression or missing dependencyReproduce locally; see testing skill

Mais skills de nvidia

fhir-basics
nvidia
Ensina aos agentes como funcionam as APIs FHIR R4, quais recursos estão disponíveis, como consultá-los com parâmetros de busca e como analisar corretamente todos os formatos de resposta…
compileiq-validate-result
nvidia
Use APÓS a conclusão de uma Pesquisa e ANTES de reivindicar qualquer aceleração ou enviar um ACF. Carrega o CSV dump_results, extrai os K melhores candidatos (objetivo único)…
changelog-audit
nvidia
Auditar o CHANGELOG.md do Warp antes de um lançamento: recuperar entradas perdidas, ordenar por impacto ao usuário, refinar a linguagem das entradas, ajustar quebras de linha e (no modo de branch de lançamento) incrementar comparação…
dgx-diagnose
nvidia
Diagnostique problemas comuns do DGX Station GB300 — falhas de CUDA, direcionamento incorreto de GPU, bugs de contêiner vLLM/SGLang, problemas de estado MIG, erros de NVLink/Fabric Manager,…
aicr-managing-openvex
nvidia
Use when adding, updating, or removing CVE/GHSA suppressions in `.openvex.json` — the OpenVEX document consumed by the daily image vulnerability scan workflow.…
aicr-creating-slide-decks
nvidia
Use ao criar um deck de slides HTML autocontido ou um ponto visual de discussão para um conceito técnico ou fluxo de trabalho (ex.: demos/*.html) — exibido em tela cheia ou…
aicr-creating-guided-demos
nvidia
Estrutura um script de demonstração guiada interativa (demos/*.sh), ao vivo ou no ritmo do usuário, com o padrão Frame → Tell → Show → Close. Aciona em "demo script", "guided…
aicr-analyzing-snapshots
nvidia
Use ao analisar um arquivo YAML de snapshot AICR, revisando o estado do cluster, comparando características de provedores, extraindo insights de topologia de GPU/rede, ou…