eagle3-validate

oleh nvidia

Memvalidasi bahwa proses pipeline EAGLE3 selesai dengan sukses dari awal hingga akhir. Memeriksa bahwa semua 4 langkah menghasilkan artefak yang diharapkan, memverifikasi tingkat penerimaan memenuhi…

npx skills add https://github.com/nvidia/model-optimizer --skill eagle3-validate

EAGLE3 Pipeline Validation

Verify that an EAGLE3 pipeline run completed successfully and meets quality criteria.

Step 0 — Identify the experiment

Find the most recent experiment directory (or ask the user for the path):

ls -td experiments/cicd/cicd_* | head -5

Each experiment directory has one subdirectory per task (numbered 0–3), each containing a log file whose name varies by launch mode (Slurm: sbatch_*.out, local Docker: *.log).

Step 1 — Check task outcomes

Match the log files generally and read the tail of each:

find experiments/<exp_id>/ -type f \( -name '*.out' -o -name '*.log' \) | sort | while read -r f; do
  echo "=== $f ==="; tail -50 "$f"; echo
done

All 4 tasks must complete without error. Look for:

  • exit code: 0 or no error — success
  • DUE TO TIME LIMIT — timeout
  • FAILED / signal / exception traceback — failure

If any task failed, suggest running /eagle3-triage instead.

Step 2 — Verify artifacts exist

Check each step produced the expected output (artifacts live on the cluster at /scratchspace/). Confirm via log messages:

StepExpected log evidenceArtifact
task_0"Saved N samples" or progress bar completing/scratchspace/data/*.jsonl
task_1"Successfully processed N conversations"/scratchspace/offline_hidden_states/*.pt
task_2Training loss decreasing, "export complete"/scratchspace/eagle3/model.safetensors, /scratchspace/export/
task_3Average Acceptance Length ... ratio: X.XXJSON result files

Step 3 — Check acceptance rate

In the task_3 log, find:

Average Acceptance Length {'accept': X, 'count': Y, 'ratio': Z.ZZ}

The ratio field is the acceptance rate (AR).

CriterionThresholdStatus
AR (MT-Bench)>= 2.1PASS / FAIL

If the log shows AR ... < lower bound, the run already triggered a threshold failure (exit code 1).

Step 4 — Check training quality

In the task_2 log look for:

  • Final training loss — should be decreasing, not NaN
  • AR validation during training (if training.ar_validate_steps was set)
  • Number of training steps — confirms full training duration

Step 5 — Produce validation report

## EAGLE3 Pipeline Validation Report

**Experiment:** <exp_dir>
**Model:** <model_name>
**Date:** <date>
**Pipeline config:** <yaml_path>

### Step Status
| Step | Task | Status | Notes |
|------|------|--------|-------|
| 0 | Data synthesis | PASS/FAIL/TIMEOUT | N samples generated |
| 1 | Hidden state dump | PASS/FAIL | N .pt files |
| 2 | Training + export | PASS/FAIL | Final loss: X.XX |
| 3 | Benchmark | PASS/FAIL | AR: X.XX |

### Acceptance Rate
- MT-Bench AR: X.XX (threshold: >= 2.1) — PASS/FAIL

### Training Summary
- Final loss: X.XX
- Training steps: N
- AR during training: X.XX (if validated)

### Overall: PASS / FAIL
<one-line summary>

Step 6 — Suggest next steps

If PASS:

  • Record the verified result (and checkpoint path) in the team's internal triage tracker
  • This model is now a candidate to add as a launcher example in a dedicated PR

If FAIL:

  • Identify which step or metric failed
  • Suggest running /eagle3-triage for diagnosis
  • For a low AR, diagnose the specific cause from the run (training loss curve, data volume/quality, draft-head capacity, hyperparameters) and suggest fixes targeted to that scenario — low AR can have many causes, so avoid a generic checklist.

Lebih banyak skill dari nvidia

compileiq-debug
nvidia
Gunakan ketika ada yang salah: Search() menggantung, semua evaluasi mengembalikan INVALID_SCORE, skor tidak kunjung membaik, setiap konfigurasi mengembalikan angka yang sama, error ptxas…
create-github-pr
nvidia
Buat pull request GitHub menggunakan gh CLI. Gunakan saat pengguna ingin membuat PR baru, mengirimkan kode untuk ditinjau, atau membuka pull request. Kata kunci pemicu -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Memindai isu terbuka lainnya untuk menemukan isu yang mungkin juga diperbaiki atau secara tidak sengaja dirusak oleh suatu PR tertentu. Menghasilkan peluang perbaikan yang berdekatan dan risiko kontradiksi dengan file:baris…
fhir-basics
nvidia
Mengajarkan agen cara kerja API FHIR R4, sumber daya apa saja yang tersedia, cara melakukan kueri dengan parameter pencarian, dan cara mengurai semua format respons dengan benar…
compileiq-validate-result
nvidia
Gunakan SETELAH Pencarian selesai dan SEBELUM mengklaim percepatan atau mengirim ACF. Muat CSV dump_results, ekstrak kandidat top-K (tujuan tunggal)…
changelog-audit
nvidia
Audit Warp CHANGELOG.md sebelum rilis: pulihkan entri yang hilang, urutkan berdasarkan dampak pengguna, perbaiki bahasa entri, bungkus baris, dan (mode cabang rilis) naikkan bandingkan…
maintain-dynamic-plugins
nvidia
Mempertahankan pemuat plugin dinamis NeMo Relay, manifes, SDK asli Rust, protokol pekerja gRPC, SDK pekerja Python, dokumen, pengujian, dan cakupan alur kerja rilis
dgx-diagnose
nvidia
Diagnosis masalah umum DGX Station GB300 — crash CUDA, penargetan GPU yang salah, bug kontainer vLLM/SGLang, masalah status MIG, kesalahan NVLink/Fabric Manager,…