eagle3-new-model

von nvidia

Fügt ein neues Modell zur EAGLE3-Offline-Pipeline hinzu. Erzeugt eine hf_offline_eagle3.yaml-Launcher-Konfiguration für einen neuen Modell-Checkpoint und wählt den richtigen versteckten Zustand aus…

npx skills add https://github.com/nvidia/model-optimizer --skill eagle3-new-model

EAGLE3 New Model Configuration

Create tools/launcher/examples/<Org>/<Model>/hf_offline_eagle3.yaml by copying the closest existing example and adapting it. Pick a reference with the same shape as the target (dense vs MoE, similar size) from tools/launcher/examples/ — e.g. the Qwen3-8B config for a dense model.

The pipeline is a 4-task config (task_0 data synthesis → task_1 hidden-state dump → task_2 train → task_3 benchmark). The task structure, args, containers, and GPU/node sizing are all visible in the existing examples — infer them from a reference rather than hand-rolling. This file documents only the two things that are not obvious from the examples: which dump backend to pick, and the model-specific gotchas.

Choosing the task_1 hidden-state dump backend

BackendScriptWhen to use
vLLMcommon/eagle3/dump_offline_data_vllm.shDefault. Broad coverage via vLLM's native hidden-state extractor.
HFcommon/eagle3/dump_offline_data_hf.shVLMs / multimodal, custom-code models, sliding-window attention (TRT-LLM can't serve these).
TRT-LLMcommon/eagle3/dump_offline_data.shPure-text models with TRT-LLM support; pass --tp <TP> and --moe-ep <EP>.

Rule of thumb: HF if the model is a VLM or uses sliding-window attention; vLLM otherwise. TRT-LLM only when you specifically want its kernels for a supported plain-text model.

Model-specific adjustments

These are the non-obvious knobs that vary per model:

SituationWhat to change
Requires --trust-remote-codeAdd to task_0 vLLM args (before the -- separator) and to task_3 benchmark args
MoE with large expert hidden dimIncrease intermediate_size in eagle_config.json to match moe_intermediate_size
Custom tokenizer (e.g. tiktoken)Set TIKTOKEN_RS_CACHE_DIR env var in task_0 and task_1

After adapting the config, preview it with --dryrun before submitting.

Mehr Skills von nvidia

compileiq-debug
nvidia
Verwenden, wenn etwas nicht stimmt: Search() hängt, alle Evaluierungen geben INVALID_SCORE zurück, Scores verbessern sich nicht, jede Konfiguration liefert dieselbe Zahl, ptxas-Fehler…
create-github-pr
nvidia
Erstelle GitHub-Pull-Requests mit der gh CLI. Verwende, wenn der Benutzer einen neuen PR erstellen, Code zur Überprüfung einreichen oder einen Pull-Request öffnen möchte. Auslöser-Schlüsselwörter -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Scannt andere offene Issues, um solche zu finden, die ein bestimmter PR möglicherweise ebenfalls behebt oder versehentlich kaputt macht. Gibt benachbarte Fix-Möglichkeiten und Widerspruchsrisiken mit Datei:Zeile… aus.
fhir-basics
nvidia
Bringt Agenten bei, wie FHIR R4 APIs funktionieren, welche Ressourcen verfügbar sind, wie man sie mit Suchparametern abfragt und wie man alle Antwortformate korrekt parst…
compileiq-validate-result
nvidia
Verwende NACH Abschluss einer Suche und VOR dem Einfordern eines Speedups oder dem Versand eines ACF. Lädt die dump_results CSV, extrahiert Top-K-Kandidaten (Einzelziel)...
changelog-audit
nvidia
Auditiere die CHANGELOG.md vor einem Release: stelle verlorene Einträge wieder her, sortiere nach Benutzerauswirkung, verfeinere die Sprache der Einträge, führe Zeilenumbrüche durch und (im Release-Branch-Modus) erhöhe die Vergleichsnummer…
maintain-dynamic-plugins
nvidia
Verwalte NeMo Relay dynamische Plugin-Lader, Manifeste, Rust native SDKs, gRPC Worker-Protokoll, Python Worker-SDK, Dokumentation, Tests und Abdeckung des Release-Workflows
dgx-diagnose
nvidia
Diagnostizieren Sie häufige DGX Station GB300-Probleme – CUDA-Abstürze, falsche GPU-Zuweisung, vLLM/SGLang-Container-Fehler, MIG-Status-Probleme, NVLink/Fabric-Manager-Fehler,…