mcore-split-pr

작성자: nvidia

PR을 여러 개로 분할하여 필요한 CODEOWNERS 리뷰어 그룹 수를 줄입니다.

npx skills add https://github.com/nvidia/megatron-lm --skill mcore-split-pr

Split PR by CODEOWNERS Groups

Split a large pull request into multiple smaller PRs, where each PR touches the fewest possible CODEOWNERS reviewer groups. The goal is to reduce review burden: a PR that only touches megatron/core/ needs only the core reviewers, while a PR that also touches examples/, tools/, and megatron/training/ pulls in many additional groups.

Answer-First Constraints

For split-planning questions, lead with these constraints before the full workflow:

  • Minimize CODEOWNERS reviewer groups per PR, but each resulting PR must still be independently mergeable and reviewable.
  • Tests travel with the production code they validate; do not split tests into a separate PR just to reduce reviewer groups.
  • If PR B depends on symbols renamed in PR A, call out the dependency and put backward-compatible aliases, re-exports, or shims in PR A when needed.
  • GitHub's standard stacked-PR flow — push each branch to the upstream repo and base each PR on the previous branch — does not work here: contributors cannot push branches to NVIDIA/Megatron-LM, and a PR's base must be an upstream branch. The only upstream refs containing a fork PR's commits are the pull-request/<N> mirrors that copy-pr-bot creates, so stacking goes through them.
  • Create every PR with base main; the pull-request/<N> mirror refs do not exist until a vetter comments /ok to test <head-sha> (copy-pr-bot). Once the mirror exists, stack a dependent PR with gh pr edit <child> --base pull-request/<base PR number>.
  • Never merge a PR while its base is a pull-request/* ref: the squash lands in the bot's scratch ref, not main, and the PR ends up MERGED and unreopenable. Retarget to main first.
  • Wait for user approval before execution.
  • Execution creates draft PRs from the right base, applies file-scoped diffs with git diff upstream/main..<source-branch> -- <paths> | git apply, pushes to the user's fork, and never pushes directly to upstream.

Workflow

1. Analyze the PR

  1. Fetch the PR details: gh pr view <number> --repo NVIDIA/Megatron-LM --json title,body,headRefName,author and gh pr diff <number> --repo NVIDIA/Megatron-LM --stat. Also determine the current GitHub user with gh api user --jq .login.
  2. Parse .github/CODEOWNERS to build a mapping from file path patterns to owner groups.
  3. For each changed file in the PR, determine which CODEOWNERS groups would be required to review it.
  4. Build a summary table grouped by CODEOWNERS group, showing which files pull in which groups.
  5. Count the total number of distinct reviewer groups the PR currently requires.

2. Propose a split that minimizes reviewer groups per PR

The primary optimization goal: minimize the number of CODEOWNERS reviewer groups required for each resulting PR.

Strategy:

  1. Cluster files by their CODEOWNERS groups. Files owned by the same set of groups naturally belong together.
  2. Identify the largest cluster — this becomes the first (and usually largest) PR.
  3. Remaining files form one or more additional PRs, each ideally requiring only one or two reviewer groups.
  4. If a split creates a dependency (e.g., PR B uses symbols renamed in PR A), the dependent PR must be merged after the first. Note this explicitly.
  5. Each PR must be independently mergeable to main — no broken imports, no missing symbols. Backward-compatible aliases and re-export stubs in the first PR can make this possible.

Present the proposed split as a table:

  • PR name/description
  • Files included
  • CODEOWNERS groups required
  • Dependencies on other PRs (if any)

Wait for user approval before proceeding.

3. Execute the split (after user approval)

For each new PR:

  1. Create a new branch from the appropriate local base (main, or a dependency PR's branch).
  2. Extract the relevant changes: git diff upstream/main..<source-branch> -- <file paths> | git apply.
  3. Stage, commit with a clear message, and push to the user's fork.
  4. Create the PR as a draft with base main (per repo contributing guidelines). Retarget dependent PRs to pull-request/<base PR number> only after a vetter's /ok to test has created that mirror ref.
  5. If the original PR needs to be narrowed in scope, confirm with the user before force-pushing.
  6. Report all PR URLs when done.

Important guidelines

  • Always create PRs as drafts and push to the user's fork, never directly to upstream.
  • Backward-compatible changes (aliases, re-exports, deprecation shims) should go in the first PR so subsequent PRs can depend on them.
  • Test files should go with the production code they test, not in a separate PR.
  • Prefer a single clean commit per split PR over replaying the original commit history.
  • If a file is hard to categorize (e.g., it touches two groups), ask the user which PR it should go in.
  • If the current GitHub user is not the author of the original PR, each new PR's description must explicitly credit the original author (e.g., "Original changes by @ in #").

nvidia의 다른 스킬

compileiq-debug
nvidia
무언가 잘못되었을 때 사용: Search()가 멈추거나, 모든 평가가 INVALID_SCORE를 반환하거나, 점수가 개선되지 않거나, 모든 설정이 동일한 숫자를 반환하거나, ptxas 오류 등이 발생할 때
create-github-pr
nvidia
gh CLI를 사용하여 GitHub 풀 리퀘스트를 생성합니다. 사용자가 새 PR을 만들거나, 코드 리뷰를 제출하거나, 풀 리퀘스트를 열고자 할 때 사용합니다. 트리거 키워드 -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
다른 열린 이슈들을 스캔하여 주어진 PR이 함께 수정하거나 실수로 망가뜨릴 수 있는 이슈를 찾습니다. 인접 수정 기회와 모순 위험을 file:line…과 함께 출력합니다.
fhir-basics
nvidia
에이전트에게 FHIR R4 API의 작동 방식, 사용 가능한 리소스, 검색 매개변수를 사용한 쿼리 방법, 모든 응답 형식을 올바르게 파싱하는 방법을 가르칩니다…
compileiq-validate-result
nvidia
검색이 완료된 후, 속도 향상을 청구하거나 ACF를 발송하기 전에 사용합니다. dump_results CSV를 로드하고, 상위 K개 후보(단일 목표)를 추출합니다…
changelog-audit
nvidia
릴리스 전에 Warp CHANGELOG.md를 감사합니다: 누락된 항목 복구, 사용자 영향별 정렬, 항목 언어 다듬기, 줄 바꿈, (릴리스 브랜치 모드) 비교 업데이트…
maintain-dynamic-plugins
nvidia
NeMo Relay 동적 플러그인 로더, 매니페스트, Rust 네이티브 SDK, gRPC 워커 프로토콜, Python 워커 SDK, 문서, 테스트 및 릴리스 워크플로 커버리지를 유지 관리합니다.
dgx-diagnose
nvidia
일반적인 DGX Station GB300 문제 진단 — CUDA 충돌, 잘못된 GPU 타겟팅, vLLM/SGLang 컨테이너 버그, MIG 상태 문제, NVLink/Fabric Manager 오류,…