cosmos3-setup

bởi nvidia

Hướng dẫn người dùng qua quá trình cài đặt Cosmos3, thiết lập môi trường, tải xuống checkpoint và xác minh. Sử dụng khi người dùng hỏi "cách cài đặt cosmos3", "cách…

npx skills add https://github.com/nvidia/cosmos-framework --skill cosmos3-setup

Cosmos3 Setup

When to use this skill

  • Use when a user wants to install Cosmos3 or set up a development environment
  • Use when a user asks about system requirements, CUDA versions, or GPU compatibility
  • Use when a user needs to download model checkpoints or configure HuggingFace auth
  • Use when a user wants to run Cosmos3 inside a Docker container or NGC container
  • For errors during setup, hand off to the cosmos3-env-troubleshoot skill

Path convention

All paths below are relative to the cosmos3 package root (../../../ from this skill file). All uv run / python commands should also be run from there.

Where to find answers

The canonical setup reference is docs/setup.md. The README (README.md § Setup) has the shortest quickstart.

User questionGo to
What are the system requirements?docs/setup.md § System Requirements
How do I install with uv? (sync, pip venv, pip system)docs/setup.md § Virtual Environment
How do I install with Docker?docs/setup.md § Docker Container
Custom torch/CUDA versions or attention backends?docs/setup.md § Advanced
Which CUDA version? (cu130 vs cu128)docs/setup.md § CUDA Variants, docs/faq.md § Which CUDA version?
How do I download checkpoints?docs/setup.md § Downloading Base Checkpoints
NGC container issues?docs/setup.md § PyTorch Import Issue
Any installation error../cosmos3-env-troubleshoot/SKILL.md

Setup steps at a glance

  1. Clone the repository and cd into the project root (the directory containing pyproject.toml)
  2. System deps: sudo apt-get install -y --no-install-recommends curl ffmpeg git-lfs libx11-dev tree wget
  3. Install uv: curl -LsSf https://astral.sh/uv/install.sh | sh && source $HOME/.local/bin/env
  4. Install package: uv sync --all-extras --group=cu130-train && source .venv/bin/activate && export LD_LIBRARY_PATH= (use cu128-train on older drivers; the inference-only cu130 / cu128 groups omit the training extras)
  5. Checkpoints: auto-downloaded during inference; requires HuggingFace auth (see docs)
  6. Verify: uv run --all-extras --group=cu130-train python -c "import cosmos_framework; print('ok')"

Things not obvious from the docs

  • NGC container caveat: you must run export LD_LIBRARY_PATH='' before any Python imports when inside an NGC PyTorch container. Easy to miss.
  • CUDA version alignment: the major CUDA version from nvidia-smi must match torch.version.cuda. Mismatches cause cryptic shared-library errors.
  • HF_HOME: controls where checkpoints are cached (default: ~/.cache/huggingface). Set this if disk space is tight or you want a shared cache.
  • Conflicting env vars: stale HF_TOKEN or HUGGING_FACE_HUB_TOKEN env vars can silently override CLI auth. Check with printenv | grep HF_.

Related skills

SkillWhen to use
../cosmos3-inference/SKILL.mdRunning inference after setup is complete
../cosmos3-codebase-nav/SKILL.mdFinding files, parameters, and configs in code
../cosmos3-env-troubleshoot/SKILL.mdDebugging environment and runtime errors

Thêm skills từ nvidia

compileiq-debug
nvidia
Sử dụng khi có điều gì đó không ổn: Search() bị treo, tất cả các đánh giá đều trả về INVALID_SCORE, điểm số không cải thiện, mọi cấu hình đều trả về cùng một số, lỗi ptxas…
create-github-pr
nvidia
Tạo pull request GitHub bằng cách sử dụng gh CLI. Sử dụng khi người dùng muốn tạo PR mới, gửi mã để xem xét, hoặc mở pull request. Từ khóa kích hoạt -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Quét các vấn đề đang mở khác để tìm những vấn đề mà một PR nhất định có thể sửa hoặc vô tình làm hỏng. Đưa ra các cơ hội sửa lỗi liền kề và rủi ro mâu thuẫn với file:dòng…
fhir-basics
nvidia
Dạy các tác nhân cách hoạt động của API FHIR R4, những tài nguyên có sẵn, cách truy vấn chúng với tham số tìm kiếm, và cách phân tích chính xác tất cả các định dạng phản hồi…
compileiq-validate-result
nvidia
Sử dụng SAU KHI tìm kiếm hoàn tất và TRƯỚC KHI yêu cầu tăng tốc hoặc gửi ACF. Tải tệp CSV dump_results, trích xuất các ứng viên top-K (đơn mục tiêu)…
changelog-audit
nvidia
Kiểm tra Warp CHANGELOG.md trước khi phát hành: khôi phục các mục bị mất, sắp xếp theo tác động người dùng, tinh chỉnh ngôn ngữ mục, xuống dòng và (chế độ nhánh phát hành) so sánh bump…
maintain-dynamic-plugins
nvidia
Duy trì các bộ nạp plugin động NeMo Relay, tệp kê khai, SDK gốc Rust, giao thức worker gRPC, SDK worker Python, tài liệu, kiểm thử và phạm vi quy trình phát hành
dgx-diagnose
nvidia
Chẩn đoán các sự cố thường gặp của DGX Station GB300 — lỗi CUDA, nhắm sai GPU, lỗi container vLLM/SGLang, vấn đề trạng thái MIG, lỗi NVLink/Fabric Manager,…