deployment

작성자: nvidia

양자화 또는 비양자화 LLM 체크포인트를 vLLM, SGLang 또는 TRT-LLM을 사용하여 OpenAI 호환 API 엔드포인트로 제공합니다. 사용자가 "모델 배포", "서빙…"이라고 말할 때 사용하세요.

npx skills add https://github.com/nvidia/model-optimizer --skill deployment

Deployment Skill

Serve a model checkpoint as an OpenAI-compatible inference endpoint. Supports vLLM, SGLang, and TRT-LLM (including AutoDeploy).

Quick Start

Prefer $SKILL_DIR/scripts/deploy.sh for standard local deployments — it handles quant detection, health checks, and server lifecycle. Use the raw framework commands in Step 4 when you need flags the script doesn't support, or for remote deployment.

# Start vLLM server with a ModelOpt checkpoint
"$SKILL_DIR/scripts/deploy.sh" start --model ./qwen3-0.6b-fp8

# Start with SGLang and tensor parallelism
"$SKILL_DIR/scripts/deploy.sh" start --model ./llama-70b-nvfp4 --framework sglang --tp 4

# Start from HuggingFace hub
"$SKILL_DIR/scripts/deploy.sh" start --model nvidia/Llama-3.1-8B-Instruct-FP8

# Test the API
"$SKILL_DIR/scripts/deploy.sh" test

# Check status
"$SKILL_DIR/scripts/deploy.sh" status

# Stop
"$SKILL_DIR/scripts/deploy.sh" stop

The script handles: GPU detection, quantization flag auto-detection (FP8 vs FP4), server lifecycle (start/stop/restart/status), health check polling, and API testing.

Decision Flow

0. Check workspace (multi-user / Slack bot)

If MODELOPT_WORKSPACE_ROOT is set, use the common skill's workspace-management.md. Before creating a new workspace, check the current session for existing model workspaces — especially if deploying a checkpoint from a prior PTQ run:

ls "$MODELOPT_WORKSPACE_ROOT/<session_id>/" 2>/dev/null

If the user says "deploy the model I just quantized" or references a previous PTQ, find the matching workspace and cd into it. The checkpoint should be in that workspace's output directory.

1. Identify the checkpoint

Determine what the user wants to deploy:

  • Local quantized checkpoint (from ptq skill or manual export): look for hf_quant_config.json in the directory. If coming from a prior PTQ run in the same workspace, check common output locations: output/, outputs/, exported_model/, or the --export_path used in the PTQ command.
  • HuggingFace model hub (e.g., nvidia/Llama-3.1-8B-Instruct-FP8): use directly
  • Unquantized model: deploy as-is (BF16) or suggest quantizing first with the ptq skill

Note: This skill expects HF-format checkpoints (from PTQ with --export_fmt hf). TRT-LLM format checkpoints should be deployed directly with TRT-LLM — see references/trtllm.md.

Check the quantization format if applicable:

cat <checkpoint_path>/hf_quant_config.json 2>/dev/null || echo "No hf_quant_config.json"

If not found, also check config.json for a quantization_config section with quant_method: "modelopt". If neither exists, the checkpoint is unquantized.

2. Choose the framework

If the user hasn't specified a framework, recommend based on this priority:

SituationRecommendedWhy
General usevLLMWidest ecosystem, easy setup, OpenAI-compatible
Best SGLang model supportSGLangStrong DeepSeek/Llama 4 support
Maximum optimizationTRT-LLMBest throughput via engine compilation
Mixed-precision / AutoQuantTRT-LLM AutoDeployOnly option for AutoQuant checkpoints

Check the support matrix in references/support-matrix.md to confirm the model + format + framework combination is supported.

3. Check the environment

Use the common skill's environment-setup.md for GPU detection, local vs remote, and SLURM/Docker/bare metal detection. After completing it you should know: GPU model/count, local or remote, and execution environment.

Then check the deployment framework is installed:

python -c "import vllm; print(f'vLLM {vllm.__version__}')" 2>/dev/null || echo "vLLM not installed"
python -c "import sglang; print(f'SGLang {sglang.__version__}')" 2>/dev/null || echo "SGLang not installed"
python -c "import tensorrt_llm; print(f'TRT-LLM {tensorrt_llm.__version__}')" 2>/dev/null || echo "TRT-LLM not installed"

If not installed, consult references/setup.md.

GPU memory estimate (to determine tensor parallelism):

  • BF16: params × 2 bytes (8B ≈ 16 GB)
  • FP8: params × 1 byte (8B ≈ 8 GB)
  • FP4: params × 0.5 bytes (8B ≈ 4 GB)
  • Add ~2-4 GB for KV cache and framework overhead

If the model exceeds single GPU memory, use tensor parallelism (-tp <num_gpus>).

4. Deploy

Read the framework-specific reference for detailed instructions:

FrameworkReference file
vLLMreferences/vllm.md
SGLangreferences/sglang.md
TRT-LLMreferences/trtllm.md

Quick-start commands (for common cases):

vLLM

# Serve as OpenAI-compatible endpoint
python -m vllm.entrypoints.openai.api_server \
    --model <checkpoint_path> \
    --quantization modelopt \
    --tensor-parallel-size <num_gpus> \
    --host 0.0.0.0 --port 8000

For NVFP4 checkpoints, use --quantization modelopt_fp4.

NVFP4 on Blackwell B300/GB300 (sm_103) needs a CUDA-13 image. From v0.20.0 on, release tags are CUDA-13 unsuffixed (e.g. vllm/vllm-openai:v0.26.0) with -cu129 the CUDA-12 opt-out; v0.19.x and earlier were the other way round (-cu130 = CUDA 13), and -cu130 no longer exists after v0.20.0. Don't trust the tag name — select a tag reporting CUDA_VERSION >= 13 in the config blob of your platform's child manifest (arm64 Grace/GB300, amd64 x86); TORCH_CUDA_ARCH_LIST differs between the two. A cu12 build has no sm_103 FP4 kernel, so vLLM loads the checkpoint then dies at engine init with CUDA error: no kernel image is available for execution on the device (affects the flashinfer and cutlass NVFP4 backends; marlin separately fails on non-64-divisible layer dims). Cross-check via recipes.vllm.ai/<org>/<model>?hardware=b300 (JS-rendered — fetch the raw markdown at github.com/vllm-project/recipes/blob/main/<org>/<model>.md). For multimodal models on sm_103, also pass --mm-encoder-attn-backend TRITON_ATTN (the default CuTe ViT flash-attn asserts "Only SM 10.x and 11.x").

SGLang

python -m sglang.launch_server \
    --model-path <checkpoint_path> \
    --quantization modelopt \
    --tp <num_gpus> \
    --host 0.0.0.0 --port 8000

For NVFP4 checkpoints, use --quantization modelopt_fp4.

Cross-check SGLang launch flags via the SGLang cookbook (the SGLang analog of recipes.vllm.ai): docs.sglang.io/cookbook/<category>/<org>/<model> (e.g. .../autoregressive/DeepSeek/DeepSeek-V4) — authoritative for parallelism, MoE backends, strategy flags, Docker image, and min version. Select the variant via the URL fragment #hw=...&variant=...&quant=...&strategy=...&nodes=.... The page is JS-rendered — fetch the raw markdown at raw.githubusercontent.com/sgl-project/sglang/main/docs_new/cookbook/<category>/<org>/<model>.mdx. SM120 (RTX PRO 6000) needs the lmsysorg/sglang:dev nightly (:latest lacks SM120). See references/sglang.md for the full backend/flag matrix.

TRT-LLM (direct)

from tensorrt_llm import LLM, SamplingParams
llm = LLM(model="<checkpoint_path>")
outputs = llm.generate(["Hello, my name is"], SamplingParams(temperature=0.8, top_p=0.95))

TRT-LLM AutoDeploy

For AutoQuant or mixed-precision checkpoints, see references/trtllm.md.

5. Verify the deployment

After the server starts, verify it's healthy:

# Health check
curl -s http://localhost:8000/health

# List models
curl -s http://localhost:8000/v1/models | python -m json.tool

# Test generation
curl -s http://localhost:8000/v1/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "<model_name>",
        "prompt": "The capital of France is",
        "max_tokens": 32
    }' | python -m json.tool

All checks must pass before reporting success to the user.

5b. Benchmark throughput/latency (optional)

If the user asks to benchmark, measure throughput/latency, or compare precisions, use AIPerf (Apache-2.0, OpenAI-compatible client benchmark). See references/benchmarking.md for install, the pre-benchmark coherence gate, the aiperf profile flags (notably --extra-inputs ignore_eos:true), suggested token shapes, and how to read profile_export_aiperf.json.

6. Remote deployment (SSH/SLURM)

If a cluster config exists (~/.config/modelopt/clusters.yaml, .agents/clusters.yaml, or .claude/clusters.yaml), or the user mentions running on a remote machine:

  1. Check container registry auth — before submitting any SLURM job with a container image, verify credentials exist on the cluster per the common skill's slurm-setup.md section 6. If credentials are missing for the image's registry, ask the user to fix auth or switch to an image on an authenticated registry (e.g., NGC). Do not submit until auth is confirmed.

  2. Source remote utilities: Load the common skill, then resolve remote_exec.sh from that skill's root.

    source "<common-skill-dir>/remote_exec.sh"
    remote_load_cluster
    remote_check_ssh
    remote_detect_env
    
  3. Sync the checkpoint (only if it was produced locally):

    If the checkpoint path is a remote/absolute path (e.g., from a prior PTQ run on the cluster), skip sync — it's already there. Verify with remote_run "ls <checkpoint_path>/config.json". Only sync if the checkpoint is local:

    remote_sync_to <local_checkpoint_path> <session_id>/<model>/checkpoints/
    
  4. Deploy based on remote environment:

    • SLURM — see the common skill's slurm-setup.md for job script templates (container setup, account/partition discovery). The server command inside the container is the same as Step 4 (e.g., python -m vllm.entrypoints.openai.api_server --model <path> --quantization modelopt). After submitting, register the job and set up monitoring per the monitor skill. Get the node hostname from squeue -j $JOBID -o %N.

    • Bare metal / Docker — use remote_run to start the server directly:

      remote_run "nohup python -m vllm.entrypoints.openai.api_server --model <path> --port 8000 > deploy.log 2>&1 &"
      
  5. Verify remotely:

    remote_run "curl -s http://localhost:8000/health"
    remote_run "curl -s http://localhost:8000/v1/models"
    
  6. Report the endpoint — include the remote hostname and port so the user can connect (e.g., http://<node_hostname>:8000). For SLURM, note that the port is only reachable from within the cluster network.

For NEL-managed deployment (evaluation with self-deployment), use the evaluation skill instead — NEL handles SLURM container deployment, health checks, and teardown automatically.

Error Handling

ErrorCauseFix
CUDA error: an illegal memory access on an NVFP4 MoE (trtllm_fused_moe_dev_kernel.cu, deepgemm)Fused-MoE FP4 kernel fault on long-context loads; engine dies, all requests 500Try VLLM_USE_FLASHINFER_MOE_FP4=1 + VLLM_FLASHINFER_MOE_BACKEND=throughput. Not always a fix — on DeepSeek-V4 it moved the fault from TRT-LLM to DeepGEMM. It also makes quantized vs baseline not kernel-matched; record that when reporting deltas.
CUDA out of memoryModel too large for GPU(s)Increase --tensor-parallel-size or use a smaller model
quantization="modelopt" not recognizedvLLM/SGLang version too oldUpgrade: vLLM >= 0.10.1, SGLang >= 0.4.10
hf_quant_config.json not foundNot a ModelOpt-exported checkpointRe-export with export_hf_checkpoint(), or remove --quantization flag
Connection refused on health checkServer still startingWait 30-60s for large models; check logs for errors
modelopt_fp4 not supportedFramework doesn't support FP4 for this modelCheck support matrix in references/support-matrix.md

Unsupported Models

If the model is not in the validated support matrix (references/support-matrix.md), deployment may fail due to weight key mismatches, missing architecture mappings, or quantized/unquantized layer confusion. Read references/unsupported-models.md for the iterative debug loop: run → read error → diagnose → patch framework source → re-run. For kernel-level issues, escalate to the framework team rather than attempting fixes.

Success Criteria

  1. Server process is running and healthy (/health returns 200)
  2. Model is listed at /v1/models
  3. Test generation produces coherent output
  4. Server URL and port are reported to the user
  5. If benchmarking was requested, throughput/latency numbers are reported

nvidia의 다른 스킬

fhir-basics
nvidia
에이전트에게 FHIR R4 API의 작동 방식, 사용 가능한 리소스, 검색 매개변수를 사용한 쿼리 방법, 모든 응답 형식을 올바르게 파싱하는 방법을 가르칩니다…
compileiq-validate-result
nvidia
검색이 완료된 후, 속도 향상을 청구하거나 ACF를 발송하기 전에 사용합니다. dump_results CSV를 로드하고, 상위 K개 후보(단일 목표)를 추출합니다…
changelog-audit
nvidia
릴리스 전에 Warp CHANGELOG.md를 감사합니다: 누락된 항목 복구, 사용자 영향별 정렬, 항목 언어 다듬기, 줄 바꿈, (릴리스 브랜치 모드) 비교 업데이트…
dgx-diagnose
nvidia
일반적인 DGX Station GB300 문제 진단 — CUDA 충돌, 잘못된 GPU 타겟팅, vLLM/SGLang 컨테이너 버그, MIG 상태 문제, NVLink/Fabric Manager 오류,…
aicr-managing-openvex
nvidia
Use when adding, updating, or removing CVE/GHSA suppressions in `.openvex.json` — the OpenVEX document consumed by the daily image vulnerability scan workflow.…
aicr-creating-slide-decks
nvidia
기술 개념이나 워크플로우에 대한 독립형 HTML 슬라이드 덱 또는 시각적 발표 자료(예: demos/*.html)를 만들 때 사용하세요. 전체 화면으로 표시하거나…
aicr-creating-guided-demos
nvidia
대화형 안내 데모 스크립트(demos/*.sh)를 라이브 또는 자기 주도 방식으로 Frame → Tell → Show → Close 패턴에 따라 구조화한다. "데모 스크립트", "안내…"와 같은 표현에 반응한다.
aicr-analyzing-snapshots
nvidia
AICR 스냅샷 YAML 파일을 분석하거나, 클러스터 상태를 검토하거나, 공급자 특성을 비교하거나, GPU/네트워크 토폴로지 인사이트를 추출할 때 사용합니다...