mcore-migrate-gpt-to-hybrid

bởi nvidia

Hướng dẫn di chuyển các checkpoint của Megatron Core GPTModel, nhà cung cấp mô hình, lệnh huấn luyện và ánh xạ lớp sang HybridModel.

npx skills add https://github.com/nvidia/megatron-lm --skill mcore-migrate-gpt-to-hybrid

GPTModel to HybridModel Migration

Answer-First Migration Guidance

  • The canonical source is docs/user-guide/hybrid-model-migration.md.
  • Read the canonical document completely before answering, planning, reviewing, editing, converting, or training.
  • Keep migration behavior, commands, mappings, prerequisites, limitations, and validation in the canonical document only. Do not duplicate them in this skill.
  • This skill adds only what the canonical document does not cover: the mechanical procedure for editing an existing launch script, and the shell hazards that procedure runs into.

Workflow

  1. Pull the task artifact first: checkpoint metadata, model provider or config, training command, conversion log, diff, or failure output.
  2. Read the canonical migration document completely.
  3. Follow only the relevant document sections. Do not invent an unsupported migration path or silently change the target architecture.
  4. Validate the result proportionately, invoking the relevant repository build and testing skills when applicable.
  5. Report the outcome and link the canonical document for human readers.

Transferring an Existing Launch Script

The canonical document specifies what the migrated command must contain. This section covers how to edit a working script into it without silent breakage. Apply the edits in order.

1. Entrypoint. pretrain_gpt.pypretrain_hybrid.py. When the entrypoint comes from a shell variable or a wrapper, follow it to the real invocation.

2. Generate the pattern instead of typing it. A 96-layer model needs a 192-character pattern; hand-typing invites a silent off-by-one.

n=32                                   # source GPT layer count
blk='*-'                               # '*-' dense, '*E' every-layer MoE
pat=$(printf "$blk%.0s" $(seq $n))

With pipeline segments (seg must divide n, and the segment count must be divisible by --pipeline-model-parallel-size):

n=32; seg=4; per=$((n/seg))
b=$(printf "$blk%.0s" $(seq $per)); pat=$b
for ((i=1;i<seg;i++)); do pat="$pat|$b"; done

3. Replace --num-layers N with --hybrid-layer-pattern. Deleting --num-layers matters: leaving a stale value is only a warning, so it looks healthy while being silently overridden by the pattern-derived count.

4. Add the stack spec, replacing any GPT --spec rather than adding a second one.

5. Delete the pipeline-layout arguments the parser rejects, and repoint --save at a new directory. See the canonical document for both lists.

Bash-array scripts: quoting at the definition site is not enough

Most scripts under examples/ collect arguments in arrays and expand them unquoted:

torchrun ${DISTRIBUTED_ARGS[@]} pretrain_gpt.py ${MODEL_ARGS[@]}

Unquoted ${ARR[@]} re-runs word-splitting and pathname expansion on every element, so the pattern is globbed against the launch directory at expansion time — single-quoting it where the array is defined does not protect it:

touch 'a-b-'; ARGS=(--hybrid-layer-pattern '*-*-')
printf '[%s]\n' ${ARGS[@]}     # -> [--hybrid-layer-pattern] [a-b-]   silently corrupted
printf '[%s]\n' "${ARGS[@]}"   # -> [--hybrid-layer-pattern] [*-*-]   correct

An unmatched glob survives intact, so this passes by luck in most working directories and fails only when some file happens to match. Store the pattern in a variable and quote that array's expansion:

HYBRID_PATTERN=$(printf '*-%.0s' $(seq $NUM_LAYERS))
MODEL_ARGS=( ... --hybrid-layer-pattern "$HYBRID_PATTERN" ... )
torchrun "${DISTRIBUTED_ARGS[@]}" pretrain_hybrid.py "${MODEL_ARGS[@]}" ...

Verify the rewrite

Both checks are cheap and catch the common slips:

# 1. No rejected or stale arguments survived -- must print nothing.
grep -nE -- '--(num-layers|num-layers-per-virtual-pipeline-stage|num-virtual-stages-per-pipeline-rank|pipeline-model-parallel-layout|account-for-embedding-in-pipeline-split|account-for-loss-in-pipeline-split|hybrid-override-pattern|fim-data)\b' train_hybrid.sh

# 2. Pattern shape -- attn and mlp must each equal the source GPT layer count.
p='*-*-|*-*-'
main=${p%%/*}; main=${main//|/}
attn=${main//[^\*]/}; mlp=${main//[^-E]/}; segs=${p%%/*}; segs=${segs//[^|]/}
echo "layers=${#main} attn=${#attn} mlp=${#mlp} segments=$(( ${#segs} + 1 ))"

Then diff the migrated script against the original: it should contain the edits above and nothing else.

Expected result of an architecture-preserving transfer

A *- or *E transfer changes the layer indexing, not the model. On a measured 8-block dense run (2 GPUs, bf16, seq 4096, 100 iterations, identical seed and data), pretrain_gpt.py --num-layers 8 and pretrain_hybrid.py --hybrid-layer-pattern '*-*-*-*-*-*-*-*-' produced:

  • identical parameter counts (2,818,641,920 on both);
  • HybridModel: ... layers='*-*-*-*-*-*-*-*-' (16 layers) from the allocator;
  • steady-state throughput within 0.1% (490.7 vs 491.2 ms/iter);
  • identical loss for the first two iterations, then a zero-mean drift of |Δ| ≤ 0.08 attributable to kernel/reduction ordering.

Treat a systematic loss offset, a parameter-count difference, or a throughput gap beyond noise as a migration bug, not as expected behavior. Note that per-iteration wall clock early in a run is dominated by dataset-cache warmup, so compare steady-state iterations only.


Documentation Drift

If the implementation and migration guide disagree:

  1. Report the discrepancy before continuing.
  2. If the task authorizes a correction, update the canonical document first.
  3. Do not add a competing migration rule to this skill.

Thêm skills từ nvidia

compileiq-debug
nvidia
Sử dụng khi có điều gì đó không ổn: Search() bị treo, tất cả các đánh giá đều trả về INVALID_SCORE, điểm số không cải thiện, mọi cấu hình đều trả về cùng một số, lỗi ptxas…
create-github-pr
nvidia
Tạo pull request GitHub bằng cách sử dụng gh CLI. Sử dụng khi người dùng muốn tạo PR mới, gửi mã để xem xét, hoặc mở pull request. Từ khóa kích hoạt -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Quét các vấn đề đang mở khác để tìm những vấn đề mà một PR nhất định có thể sửa hoặc vô tình làm hỏng. Đưa ra các cơ hội sửa lỗi liền kề và rủi ro mâu thuẫn với file:dòng…
fhir-basics
nvidia
Dạy các tác nhân cách hoạt động của API FHIR R4, những tài nguyên có sẵn, cách truy vấn chúng với tham số tìm kiếm, và cách phân tích chính xác tất cả các định dạng phản hồi…
compileiq-validate-result
nvidia
Sử dụng SAU KHI tìm kiếm hoàn tất và TRƯỚC KHI yêu cầu tăng tốc hoặc gửi ACF. Tải tệp CSV dump_results, trích xuất các ứng viên top-K (đơn mục tiêu)…
changelog-audit
nvidia
Kiểm tra Warp CHANGELOG.md trước khi phát hành: khôi phục các mục bị mất, sắp xếp theo tác động người dùng, tinh chỉnh ngôn ngữ mục, xuống dòng và (chế độ nhánh phát hành) so sánh bump…
maintain-dynamic-plugins
nvidia
Duy trì các bộ nạp plugin động NeMo Relay, tệp kê khai, SDK gốc Rust, giao thức worker gRPC, SDK worker Python, tài liệu, kiểm thử và phạm vi quy trình phát hành
dgx-diagnose
nvidia
Chẩn đoán các sự cố thường gặp của DGX Station GB300 — lỗi CUDA, nhắm sai GPU, lỗi container vLLM/SGLang, vấn đề trạng thái MIG, lỗi NVLink/Fabric Manager,…