cosmos3-inference

작성자: nvidia

Cosmos3 추론 실행 방법을 사용자에게 안내합니다 — 오프라인 배치 생성, Ray 및 Gradio를 사용한 온라인 서빙, 병렬 처리 옵션, 입력 형식, 샘플링…

npx skills add https://github.com/nvidia/cosmos-framework --skill cosmos3-inference

Cosmos3 Inference

When to use this skill

  • Use when a user wants to generate images or videos with Cosmos3
  • Use when a user asks about inference parameters, input formats, or parallelism
  • Use when a user wants to set up online serving (Ray Serve, Gradio)
  • Use when a user asks about prompt engineering or upsampling
  • For environment or import errors, hand off to cosmos3-env-troubleshoot

Path convention

All paths below are relative to the cosmos3 package root (../../../ from this skill file). All uv run / python commands should also be run from there.

Where to find answers

User questionGo to
How do I run inference? (single-GPU, multi-GPU)README.md § Inference
Which model should I use? (Nano vs Super, memory, shift)README.md § Models
Which modality? (t2i, t2v, i2v, examples)README.md § Modalities
What parallelism preset? (latency vs throughput)README.md § Inference
What input fields are available? (prompt, vision_path, num_frames, ...)docs/inference.md § Sample Arguments
What are the default parameter values?cosmos_framework/inference/defaults/<model_mode>/sample_args.json (per-modality JSON)
How do I use custom defaults?docs/inference.md § Custom Defaults
How do I override a parameter? (precedence)docs/faq.md § How do I override a default parameter?
What is the shift parameter?docs/faq.md § What is the shift parameter?
How many frames can I generate? (resolution caps)docs/faq.md § How many frames can I generate?
How do I start Ray Serve / Gradio / submit requests?docs/faq.md § How do I run online inference with Ray?
How do I upsample short prompts?docs/faq.md § Prompt upsampling
How do I use the low-level API? (examples/)examples/inference.py (model API) / examples/inference_pipeline.py (pipeline API)
All CLI flagsuv run --all-extras --group=cu130 python -m cosmos_framework.scripts.inference --help

Things not obvious from the docs

  • Path resolution: relative paths in input JSON files are resolved relative to the JSON file's directory, not the working directory.
  • Seed: always pass --seed for reproducible results. Without it, a random seed is used each time.
  • Resume: interrupted runs can be resumed by re-running the same command — existing outputs are skipped automatically.
  • --keep-going: continues processing remaining samples after a per-sample failure (e.g. guardrail rejection). Used in online serving by default.
  • Unique names: every sample in a run must have a unique name field, or the script will error.

Related skills

SkillWhen to use
../cosmos3-setup/SKILL.mdInstallation and environment setup
../cosmos3-codebase-nav/SKILL.mdFinding files, parameters, and configs in code
../cosmos3-env-troubleshoot/SKILL.mdDebugging environment and runtime errors

nvidia의 다른 스킬

compileiq-debug
nvidia
무언가 잘못되었을 때 사용: Search()가 멈추거나, 모든 평가가 INVALID_SCORE를 반환하거나, 점수가 개선되지 않거나, 모든 설정이 동일한 숫자를 반환하거나, ptxas 오류 등이 발생할 때
create-github-pr
nvidia
gh CLI를 사용하여 GitHub 풀 리퀘스트를 생성합니다. 사용자가 새 PR을 만들거나, 코드 리뷰를 제출하거나, 풀 리퀘스트를 열고자 할 때 사용합니다. 트리거 키워드 -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
다른 열린 이슈들을 스캔하여 주어진 PR이 함께 수정하거나 실수로 망가뜨릴 수 있는 이슈를 찾습니다. 인접 수정 기회와 모순 위험을 file:line…과 함께 출력합니다.
fhir-basics
nvidia
에이전트에게 FHIR R4 API의 작동 방식, 사용 가능한 리소스, 검색 매개변수를 사용한 쿼리 방법, 모든 응답 형식을 올바르게 파싱하는 방법을 가르칩니다…
compileiq-validate-result
nvidia
검색이 완료된 후, 속도 향상을 청구하거나 ACF를 발송하기 전에 사용합니다. dump_results CSV를 로드하고, 상위 K개 후보(단일 목표)를 추출합니다…
changelog-audit
nvidia
릴리스 전에 Warp CHANGELOG.md를 감사합니다: 누락된 항목 복구, 사용자 영향별 정렬, 항목 언어 다듬기, 줄 바꿈, (릴리스 브랜치 모드) 비교 업데이트…
maintain-dynamic-plugins
nvidia
NeMo Relay 동적 플러그인 로더, 매니페스트, Rust 네이티브 SDK, gRPC 워커 프로토콜, Python 워커 SDK, 문서, 테스트 및 릴리스 워크플로 커버리지를 유지 관리합니다.
dgx-diagnose
nvidia
일반적인 DGX Station GB300 문제 진단 — CUDA 충돌, 잘못된 GPU 타겟팅, vLLM/SGLang 컨테이너 버그, MIG 상태 문제, NVLink/Fabric Manager 오류,…