eagle3-review-logs

作成者: nvidia

EAGLE3パイプラインの実験ログをランチャーのexperiments/ディレクトリからレビューします。4つのタスクすべての合格/不合格ステータスを要約し、根本原因による障害を診断します…

npx skills add https://github.com/nvidia/model-optimizer --skill eagle3-review-logs

Review EAGLE3 Experiment Logs

Analyze output logs from an EAGLE3 pipeline run launched via launch.py or slurm.py.

Step 0 — Find experiment logs

Locate the experiment directory. The default is experiments/ relative to the launcher root, or wherever --job-dir was pointed.

ls -td experiments/cicd/cicd_* | head -10

If no experiments exist, ask the user for the directory.

Step 1 — Read all task logs

Each experiment has one subdirectory per task (0–3). Log filenames vary by launch mode (Slurm writes sbatch_*.out, local Docker writes *.log), so match log files generally and read the tail of each in a single Bash call — errors surface at the end:

find experiments/<exp_id>/ -type f \( -name '*.out' -o -name '*.log' \) | sort | while read -r f; do
  echo "=== $f ==="; tail -200 "$f"; echo
done

Step 2 — Analyze

For each task log, check:

  • Exit / cancellation: DUE TO TIME LIMIT, FAILED, signal (e.g., signal 15)
  • Python exceptions / tracebacks: last exception is usually the root cause
  • CUDA errors: OOM, NCCL timeout
  • Slurm state: COMPLETED, FAILED, TIMEOUT, OUT_OF_MEMORY
  • Success indicators: "Saved N samples", "Successfully processed N conversations", training loss line, AR output

Step 3 — Produce report

Output a structured markdown report:

Summary

  • Overall status: PASSED / FAILED / MIXED / PARTIAL
  • Task breakdown: e.g., task_0 TIMEOUT, task_1 FAIL, task_2 skipped, task_3 skipped

Task Results

For each task (0–3):

Task N — <name>: PASS / FAIL / TIMEOUT

  • Key output: (e.g., "3277/3295 samples generated" or "Script not found")
  • Error (if failed): quoted error message, max 10 lines
  • Root cause: one-line diagnosis
  • Suggested fix: actionable step

Warnings

Non-fatal issues worth noting (near-OOM, tokenizer warnings, slow throughput).

Step 4 — Suggest next steps

Based on results:

  • If a task failed due to a known issue, suggest the fix and how to re-run from that task:

    uv run launch.py --yaml examples/<Org>/<Model>/hf_offline_eagle3.yaml \
        pipeline.task_0.skip=true \
        --yes
    
  • If the failure pattern looks new, suggest capturing it in the team's internal triage tracker, and use /eagle3-triage for a deeper diagnosis.

  • If all tasks passed, suggest running /eagle3-validate to confirm AR meets threshold.

Known benign patterns (do NOT mark as failures)

PatternExplanation
vLLM server exit code 143SIGTERM — server was killed after queries completed. Expected.
CANCELLED AT ... DUE TO TASK FAILURE after exit code: 0Slurm cleanup of worker nodes after main task succeeded.
destroy_process_group() was not calledBenign PyTorch shutdown warning.
tokenizer class ... not equal to the registered tokenizer classHarmless tokenizer mismatch warning.

nvidiaのその他のスキル

compileiq-debug
nvidia
何かがおかしいときに使用:Search()がハングする、すべての評価がINVALID_SCOREを返す、スコアが改善しない、すべての設定が同じ数値を返す、ptxasエラー…
create-github-pr
nvidia
gh CLIを使用してGitHubのプルリクエストを作成します。ユーザーが新しいPRを作成したい、コードをレビューに提出したい、またはプルリクエストを開きたい場合に使用します。トリガーキーワード -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
他のオープンなIssueをスキャンし、特定のPRが修正する可能性があるものや、誤って壊す可能性があるものを見つけます。隣接修正の機会や矛盾リスクをfile:line…と共に出力します。
fhir-basics
nvidia
エージェントにFHIR R4 APIの動作方法、利用可能なリソース、検索パラメータを使ったクエリ方法、およびすべてのレスポンス形式を正しく解析する方法を教えます…
compileiq-validate-result
nvidia
検索が完了した後、かつスピードアップの申請やACFの発送の前に使用します。dump_results CSVを読み込み、トップK候補(単一目的)を抽出します…
changelog-audit
nvidia
リリース前にWarp CHANGELOG.mdを監査:失われたエントリを復元、ユーザー影響で並べ替え、エントリの文言を洗練、行折り返し、および(リリースブランチモードで)比較をバンプ…
maintain-dynamic-plugins
nvidia
NeMo Relayの動的プラグインローダー、マニフェスト、RustネイティブSDK、gRPCワーカープロトコル、PythonワーカーSDK、ドキュメント、テスト、およびリリースワークフローのカバレッジを維持する
dgx-diagnose
nvidia
一般的なDGX Station GB300の問題(CUDAクラッシュ、誤ったGPUターゲット、vLLM/SGLangコンテナのバグ、MIG状態の問題、NVLink/Fabric Managerエラーなど)を診断します。