eagle3-validate

작성자: nvidia

EAGLE3 파이프라인 실행이 종단 간 성공적으로 완료되었는지 검증합니다. 4단계 모두에서 예상된 아티팩트가 생성되었는지 확인하고, 수용률이 기준을 충족하는지 검증합니다…

npx skills add https://github.com/nvidia/model-optimizer --skill eagle3-validate

EAGLE3 Pipeline Validation

Verify that an EAGLE3 pipeline run completed successfully and meets quality criteria.

Step 0 — Identify the experiment

Find the most recent experiment directory (or ask the user for the path):

ls -td experiments/cicd/cicd_* | head -5

Each experiment directory has one subdirectory per task (numbered 0–3), each containing a log file whose name varies by launch mode (Slurm: sbatch_*.out, local Docker: *.log).

Step 1 — Check task outcomes

Match the log files generally and read the tail of each:

find experiments/<exp_id>/ -type f \( -name '*.out' -o -name '*.log' \) | sort | while read -r f; do
  echo "=== $f ==="; tail -50 "$f"; echo
done

All 4 tasks must complete without error. Look for:

  • exit code: 0 or no error — success
  • DUE TO TIME LIMIT — timeout
  • FAILED / signal / exception traceback — failure

If any task failed, suggest running /eagle3-triage instead.

Step 2 — Verify artifacts exist

Check each step produced the expected output (artifacts live on the cluster at /scratchspace/). Confirm via log messages:

StepExpected log evidenceArtifact
task_0"Saved N samples" or progress bar completing/scratchspace/data/*.jsonl
task_1"Successfully processed N conversations"/scratchspace/offline_hidden_states/*.pt
task_2Training loss decreasing, "export complete"/scratchspace/eagle3/model.safetensors, /scratchspace/export/
task_3Average Acceptance Length ... ratio: X.XXJSON result files

Step 3 — Check acceptance rate

In the task_3 log, find:

Average Acceptance Length {'accept': X, 'count': Y, 'ratio': Z.ZZ}

The ratio field is the acceptance rate (AR).

CriterionThresholdStatus
AR (MT-Bench)>= 2.1PASS / FAIL

If the log shows AR ... < lower bound, the run already triggered a threshold failure (exit code 1).

Step 4 — Check training quality

In the task_2 log look for:

  • Final training loss — should be decreasing, not NaN
  • AR validation during training (if training.ar_validate_steps was set)
  • Number of training steps — confirms full training duration

Step 5 — Produce validation report

## EAGLE3 Pipeline Validation Report

**Experiment:** <exp_dir>
**Model:** <model_name>
**Date:** <date>
**Pipeline config:** <yaml_path>

### Step Status
| Step | Task | Status | Notes |
|------|------|--------|-------|
| 0 | Data synthesis | PASS/FAIL/TIMEOUT | N samples generated |
| 1 | Hidden state dump | PASS/FAIL | N .pt files |
| 2 | Training + export | PASS/FAIL | Final loss: X.XX |
| 3 | Benchmark | PASS/FAIL | AR: X.XX |

### Acceptance Rate
- MT-Bench AR: X.XX (threshold: >= 2.1) — PASS/FAIL

### Training Summary
- Final loss: X.XX
- Training steps: N
- AR during training: X.XX (if validated)

### Overall: PASS / FAIL
<one-line summary>

Step 6 — Suggest next steps

If PASS:

  • Record the verified result (and checkpoint path) in the team's internal triage tracker
  • This model is now a candidate to add as a launcher example in a dedicated PR

If FAIL:

  • Identify which step or metric failed
  • Suggest running /eagle3-triage for diagnosis
  • For a low AR, diagnose the specific cause from the run (training loss curve, data volume/quality, draft-head capacity, hyperparameters) and suggest fixes targeted to that scenario — low AR can have many causes, so avoid a generic checklist.

nvidia의 다른 스킬

compileiq-debug
nvidia
무언가 잘못되었을 때 사용: Search()가 멈추거나, 모든 평가가 INVALID_SCORE를 반환하거나, 점수가 개선되지 않거나, 모든 설정이 동일한 숫자를 반환하거나, ptxas 오류 등이 발생할 때
create-github-pr
nvidia
gh CLI를 사용하여 GitHub 풀 리퀘스트를 생성합니다. 사용자가 새 PR을 만들거나, 코드 리뷰를 제출하거나, 풀 리퀘스트를 열고자 할 때 사용합니다. 트리거 키워드 -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
다른 열린 이슈들을 스캔하여 주어진 PR이 함께 수정하거나 실수로 망가뜨릴 수 있는 이슈를 찾습니다. 인접 수정 기회와 모순 위험을 file:line…과 함께 출력합니다.
fhir-basics
nvidia
에이전트에게 FHIR R4 API의 작동 방식, 사용 가능한 리소스, 검색 매개변수를 사용한 쿼리 방법, 모든 응답 형식을 올바르게 파싱하는 방법을 가르칩니다…
compileiq-validate-result
nvidia
검색이 완료된 후, 속도 향상을 청구하거나 ACF를 발송하기 전에 사용합니다. dump_results CSV를 로드하고, 상위 K개 후보(단일 목표)를 추출합니다…
changelog-audit
nvidia
릴리스 전에 Warp CHANGELOG.md를 감사합니다: 누락된 항목 복구, 사용자 영향별 정렬, 항목 언어 다듬기, 줄 바꿈, (릴리스 브랜치 모드) 비교 업데이트…
maintain-dynamic-plugins
nvidia
NeMo Relay 동적 플러그인 로더, 매니페스트, Rust 네이티브 SDK, gRPC 워커 프로토콜, Python 워커 SDK, 문서, 테스트 및 릴리스 워크플로 커버리지를 유지 관리합니다.
dgx-diagnose
nvidia
일반적인 DGX Station GB300 문제 진단 — CUDA 충돌, 잘못된 GPU 타겟팅, vLLM/SGLang 컨테이너 버그, MIG 상태 문제, NVLink/Fabric Manager 오류,…