compare-results

โดย nvidia

จัดทำแผนประเมินผลระหว่าง baseline กับ candidate มอบหมายการประเมินที่ขาดหาย เปรียบเทียบผลลัพธ์ที่ตรวจสอบแล้ว และตัดสินใจความเป็นไปได้ในการควอนไทซ์ ใช้เมื่อ...

npx skills add https://github.com/nvidia/model-optimizer --skill compare-results

Compare Results

Use this to plan and complete a baseline-vs-candidate comparison. The baseline is the reference checkpoint, and the candidate is the checkpoint whose accuracy change is being measured, typically a further quantized version of the baseline.

Workflow

  1. Establish the candidate checkpoint/run and the matching baseline. Infer the baseline from the PTQ source model/checkpoint in the workspace or config used to create the candidate. If it cannot be inferred, ask the user for the baseline checkpoint or an existing baseline invocation/run path.
  2. If a required baseline or candidate evaluation is missing, delegate to the evaluation skill to create, run, and verify it. The companion evaluation config should match benchmark versions, task configs, serving args, token limits, dataset setup, credentials, cluster, and container as closely as possible; change only the model/checkpoint and checkpoint-specific serving or quantization flags.
  3. Fetch the baseline and candidate task list, configs, score artifacts, and logs. If the user provides MLflow runs or invocation IDs, use the accessing-mlflow skill to fetch configs and artifacts.
  4. Confirm each run passed evaluation Step 9, "Verify completed evaluation run", before comparing scores. If not, validate logs, server health, judge/code-execution status, sample accounting, and reasoning parsing before computing deltas.
  5. For each task, use the canonical score field from the matching evaluation skill task recipe, recipes/tasks/<task>.md, under Score Extraction.
  6. Use the evaluation skill's references/run-validation.md to perform the External Baseline Sanity Check. Record each source URL, protocol difference, and task status before applying the candidate-delta gate. A failed baseline blocks a success verdict; correct and rerun it first. If no credible comparable reference exists, label the baseline externally unverified rather than claiming the check passed, then continue using the validated measured baseline.
  7. Compute exact deltas outside the chat context when there are multiple tasks or repeated runs.
  8. Report comparability, external baseline sanity, and quantized-feasibility verdicts before interpreting the delta as model quality. If the user did not provide an acceptance threshold, report feasibility as inconclusive instead of inventing one.

Comparability Checklist

Before treating a baseline-vs-quantized delta as a model quality result, verify the validated runs are comparable:

  1. Prompt text, system prompt, chat template, and rendered messages match.
  2. Task name, benchmark version, dataset split, container, harness, and task fragment match.
  3. Generation settings match, including temperature, top_p, top_k, max tokens, stop strings, chat-template kwargs, reasoning mode/budget, and task-specific overrides.
  4. Reasoning traces are enabled, disabled, parsed, stripped, or ignored consistently between runs.
  5. The number of evaluated and scored samples/repeats matches for each task and split.
  6. Judge-backed or simulator-backed tasks use the same judge/user model, endpoint class, prompt, and scoring config.
  7. The same accuracy metric and score field is used for both runs.
  8. Timeout policies, effective limits, and failure scoring/exclusions match. Apply Timeout and Output-Limit Accounting in the evaluation skill's references/run-validation.md to both runs; matching limits alone cannot rule out serving-speed effects on scores.
  9. Baseline precision matches the gate. A <1pp vs BF16 gate requires a true full-precision (BF16) baseline. Many models ship natively quantized (e.g. INT4 W4A16 or block-wise FP8) with no BF16 release — a quant-to-quant comparison against the released precision (e.g. INT4 vs NVFP4, as for Kimi-K2.6) is still a valid result. State which precision the baseline is and apply only an acceptance criterion explicitly defined for that precision. If the requested gate is relative to BF16 and no BF16 baseline is available, report that gate as inconclusive; do not reinterpret it as an FP8/INT4 gate.

For SciCode, keep num_repeats: 1 and require at least 8 runs per side, comparing the two means — see the evaluation skill's recipes/tasks/aa/scicode.md. Fewer than 8 valid runs on a side is INDETERMINATE, not a delta.

If any item differs, either rerun with matched settings or label the result as not an apples-to-apples quantization comparison.

These checks compare the baseline and candidate to each other. The external baseline check in the evaluation skill's references/run-validation.md separately tests whether the baseline's absolute score is credible; both guards must be reported.

Report Format

Include:

  • Baseline and candidate identifiers.
  • Per-task metric path, baseline score, candidate score, delta, and stderr if available.
  • Per-task external reference score, source URL, known protocol differences, percentage-point difference, and sanity status (verified, failed, or externally unverified).
  • Comparability status for prompt/template, generation settings, sample counts, reasoning handling, judge/simulator setup, and score field.
  • Per-task timeout and output-limit counts/rates for both runs, explicit denominators, telemetry coverage, effective limits, and unresolved effects on the score. Unknown accounting or unresolved infrastructure effects prevent an acceptable quantization-feasibility verdict.
  • Comparability verdict: comparable, not comparable, or inconclusive.
  • Quantization feasibility verdict: acceptable, not acceptable, or inconclusive. Never report acceptable when external baseline sanity failed. An externally unverified baseline does not block acceptable; apply the candidate-delta gate and report the missing external corroboration.

Skills เพิ่มเติมจาก nvidia

fhir-basics
nvidia
สอนให้เอเจนต์เข้าใจการทำงานของ FHIR R4 API ทรัพยากรที่มีให้ วิธีค้นหาด้วยพารามิเตอร์ค้นหา และวิธีแยกวิเคราะห์รูปแบบการตอบกลับทั้งหมดอย่างถูกต้อง…
compileiq-validate-result
nvidia
ใช้หลังจากที่การค้นหาเสร็จสิ้น และก่อนที่จะอ้างสิทธิ์การเร่งความเร็วหรือจัดส่ง ACF โหลดไฟล์ CSV dump_results แยกผู้สมัคร K อันดับแรก (วัตถุประสงค์เดียว)…
changelog-audit
nvidia
ตรวจสอบ Warp CHANGELOG.md ก่อนปล่อย: กู้คืนรายการที่สูญหาย จัดเรียงตามผลกระทบต่อผู้ใช้ ปรับปรุงภาษาในรายการ จัดบรรทัด และ (ในโหมดสาขาปล่อย) เปรียบเทียบการเพิ่มเวอร์ชัน…
dgx-diagnose
nvidia
วินิจฉัยปัญหาทั่วไปของ DGX Station GB300 — CUDA ล่ม, การกำหนดเป้าหมาย GPU ผิด, บั๊กคอนเทนเนอร์ vLLM/SGLang, ปัญหาสถานะ MIG, ข้อผิดพลาด NVLink/Fabric Manager,…
aicr-managing-openvex
nvidia
Use when adding, updating, or removing CVE/GHSA suppressions in `.openvex.json` — the OpenVEX document consumed by the daily image vulnerability scan workflow.…
aicr-creating-slide-decks
nvidia
ใช้เมื่อสร้างสไลด์เด็ค HTML แบบครบวงจรหรือจุดนำเสนอภาพสำหรับแนวคิดทางเทคนิคหรือขั้นตอนการทำงาน (เช่น demos/*.html) — แสดงแบบเต็มหน้าจอหรือ…
aicr-creating-guided-demos
nvidia
สร้างสคริปต์เดโมแบบอินเทอร์แอกทีฟพร้อมคำแนะนำ (demos/*.sh) แบบสดหรือแบบเรียนรู้ด้วยตนเอง ตามรูปแบบ Frame → Tell → Show → Close เริ่มทำงานเมื่อมีคำว่า "สคริปต์เดโม" "แบบแนะนำ…
aicr-analyzing-snapshots
nvidia
ใช้เมื่อวิเคราะห์ไฟล์ YAML สแนปช็อต AICR ตรวจสอบสถานะคลัสเตอร์ เปรียบเทียบคุณลักษณะของผู้ให้บริการ ดึงข้อมูลเชิงลึกเกี่ยวกับโทโพโลยี GPU/เครือข่าย หรือ…