cicd

por nvidia

Referência de CI/CD para NeMo-RL. Abrange a estrutura do pipeline do GitHub Actions, acionamento de CI via /ok to test e investigação de falhas de CI.

npx skills add https://github.com/NVIDIA-NeMo/RL --skill cicd

CI/CD Guide


How CI Works

NeMo-RL CI runs on GitHub Actions. Workflows live in .github/workflows/.

The main test workflow triggers on pushes to pull-request/<number> branches. These branches are created automatically by copy-pr-bot when a contributor pushes to their fork and the PR receives a trust signal:

  • The commit is GPG-signed by a maintainer, or
  • A maintainer posts /ok to test <full-sha> as a PR comment.

Triggering CI

After pushing a new commit to your PR, trigger CI with:

/ok to test <full-commit-sha>

Use git rev-parse HEAD (not the short form) to get the full SHA.

Re-triggering after new commits: each /ok to test <sha> is specific to that SHA. After pushing additional commits, post a new comment with the updated SHA.


CI Labels

Every PR must have exactly one CI:* label — the quality-check job stays red until one is attached. Labels control which test tier runs and whether a new container image is built.

LabelWhat runsContainer
CI:docsDoc tests onlyReuses main container
CI:LfastFast test subsetReuses main container
CI:L0Unit tests + docs + lintBuilds new image
CI:L1L0 + functional testsBuilds new image
CI:L2L1 + convergence testsBuilds new image
Skip CICDNothing (skips all tests)

Default on merge group / push to main: L1.

Which label to attach when opening a PR:

Changed paths / nature of changeLabel
Docs only (docs/, *.md, docstrings)CI:docs
Trivial fix, no logic changeCI:Lfast
New code, bug fix, refactorCI:L0
Changes that could affect model behaviourCI:L1
Changes that could affect convergenceCI:L2

CI Failure Investigation

# List recent workflow runs for the PR
gh run list --repo NVIDIA-NeMo/RL --branch "pull-request/<pr-number>"

# View failing run summary
gh run view <run-id> --repo NVIDIA-NeMo/RL

# Stream failing job output
gh run view <run-id> --repo NVIDIA-NeMo/RL --log-failed

Common failure patterns:

SymptomLikely causeFix
CI never startsNo trust signal or copy-pr-bot not triggeredPost /ok to test <sha>
semantic-pull-request failsPR title doesn't follow Conventional CommitsFix PR title; see contributing skill
Linting failsStyle violationRun uv run ruff check --fix . && uv run ruff format .
Unit test failureCode regression or missing dependencyReproduce locally; see testing skill

Mais skills de nvidia

compileiq-debug
nvidia
Use quando algo está errado: Search() trava, todas as avaliações retornam INVALID_SCORE, as pontuações não estão melhorando, toda configuração retorna o mesmo número, erros de ptxas…
create-github-pr
nvidia
Crie pull requests do GitHub usando a CLI gh. Use quando o usuário quiser criar um novo PR, enviar código para revisão ou abrir um pull request. Palavras-chave de acionamento -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Escaneia outras issues abertas para encontrar aquelas que um determinado PR pode também corrigir ou quebrar acidentalmente. Gera oportunidades de correção adjacentes e riscos de contradição com arquivo:linha…
fhir-basics
nvidia
Ensina aos agentes como funcionam as APIs FHIR R4, quais recursos estão disponíveis, como consultá-los com parâmetros de busca e como analisar corretamente todos os formatos de resposta…
compileiq-validate-result
nvidia
Use APÓS a conclusão de uma Pesquisa e ANTES de reivindicar qualquer aceleração ou enviar um ACF. Carrega o CSV dump_results, extrai os K melhores candidatos (objetivo único)…
changelog-audit
nvidia
Auditar o CHANGELOG.md do Warp antes de um lançamento: recuperar entradas perdidas, ordenar por impacto ao usuário, refinar a linguagem das entradas, ajustar quebras de linha e (no modo de branch de lançamento) incrementar comparação…
maintain-dynamic-plugins
nvidia
Manter carregadores de plugins dinâmicos do NeMo Relay, manifestos, SDKs nativos em Rust, protocolo de worker gRPC, SDK de worker Python, documentação, testes e cobertura do fluxo de lançamento
dgx-diagnose
nvidia
Diagnostique problemas comuns do DGX Station GB300 — falhas de CUDA, direcionamento incorreto de GPU, bugs de contêiner vLLM/SGLang, problemas de estado MIG, erros de NVLink/Fabric Manager,…