cosmos3-env-troubleshoot

작성자: nvidia

Cosmos3 환경, 설치, 런타임 오류를 진단하고 수정합니다. 사용자가 ImportError, ModuleNotFoundError, CUDA error, Docker… 문제를 겪을 때 사용합니다.

npx skills add https://github.com/nvidia/cosmos-framework --skill cosmos3-env-troubleshoot

Cosmos3 Environment Troubleshooting

When to use this skill

  • Use when a user hits an error during installation, environment setup, or first run
  • Use when a traceback mentions torch, CUDA, missing modules, or shared libraries
  • Use when Docker or container setup fails
  • Use when checkpoint downloads fail or HuggingFace auth errors appear

Path convention

All paths below are relative to this file's location (.agents/skills/cosmos3-env-troubleshoot/).

Step 1: Match against known errors

Check the error message against the table below. Each row links to the canonical fix in the docs.

Error signatureCauseFix location
ImportError: cannot import name '_functionalization' from 'torch._C'NGC container library conflict../../../docs/setup.md § PyTorch Import Issue — run export LD_LIBRARY_PATH=''
ModuleNotFoundError: No module named 'cosmos_framework'Package not installed../../../docs/setup.md § Dependency Issue — run uv sync --all-extras --group=cu130-train --reinstall
ModuleNotFoundError: No module named <other>Dependency missing../../../docs/setup.md § Dependency Issue — reinstall venv
fatal error: Python.h: No such file or directoryBroken Python / uv install../../../docs/setup.md § Python Issue — reinstall uv + venv from scratch
OSError: <lib>: cannot open shared object fileCUDA version mismatch../../../docs/setup.md § CUDA Issue — install matching cuda-toolkit-<major>
docker: Error response from daemon: unknown or invalid runtime name: nvidiaDocker nvidia runtime not configured../../../docs/setup.md § Docker Container — run sudo nvidia-ctk runtime configure --runtime=docker
HuggingFace 401 / download failuresAuth or license not accepted../../../docs/setup.md § Downloading Base Checkpoints — check HF_TOKEN, accept license agreement

Step 2: If no documented fix matches, try common remediation

Run these diagnostic commands to collect information, then attempt fixes in order:

Diagnostic commands

# System
uname -a
cat /etc/os-release | head -5

# Python
python --version
which python

# CUDA
nvidia-smi
python -c "import torch; print(f'torch={torch.__version__}, cuda={torch.version.cuda}')"

# Package
uv pip list | head -20

Remediation ladder (try in order)

  1. Clear library path: export LD_LIBRARY_PATH=''

  2. Reinstall venv: uv sync --all-extras --group=cu130-train --reinstall (or cu128-train on older drivers; drop -train only if you intentionally want the inference-only group)

  3. Reinstall uv + venv from scratch:

    curl -LsSf https://astral.sh/uv/install.sh | sh
    uv python install --reinstall
    rm -rf .venv
    uv sync --all-extras --group=cu130-train --reinstall
    source .venv/bin/activate
    
  4. Check CUDA version alignment: the major CUDA version from nvidia-smi must match torch.version.cuda

  5. Try Docker: if the host environment is too broken, fall back to the Docker container (see ../../../docs/setup.md)

Step 3: If still unresolved, generate a bug report

If none of the above resolves the issue, collect environment information and present the user with a pre-filled bug report they can submit as a GitHub issue.

Fill in the template below by running the diagnostic commands and inserting the results:

## Environment

- **OS**: <output of `uname -a`>
- **Python**: <output of `python --version`>
- **CUDA (system)**: <output of `nvidia-smi` — first line with driver/CUDA version>
- **CUDA (torch)**: <output of `python -c "import torch; print(torch.version.cuda)">`>
- **torch version**: <output of `python -c "import torch; print(torch.__version__)">`>
- **cosmos_framework version**: <output of `python -c "import cosmos_framework; print(cosmos_framework.__version__)"` or "not installed">
- **Installation method**: <uv sync / uv pip / Docker / NGC container>

## Error

```
<full traceback>
```

## What was tried

1. <list each remediation step attempted and its result>

## Additional context

<any other relevant details — multi-GPU setup, custom CUDA install, etc.>

nvidia의 다른 스킬

fhir-basics
nvidia
에이전트에게 FHIR R4 API의 작동 방식, 사용 가능한 리소스, 검색 매개변수를 사용한 쿼리 방법, 모든 응답 형식을 올바르게 파싱하는 방법을 가르칩니다…
compileiq-validate-result
nvidia
검색이 완료된 후, 속도 향상을 청구하거나 ACF를 발송하기 전에 사용합니다. dump_results CSV를 로드하고, 상위 K개 후보(단일 목표)를 추출합니다…
changelog-audit
nvidia
릴리스 전에 Warp CHANGELOG.md를 감사합니다: 누락된 항목 복구, 사용자 영향별 정렬, 항목 언어 다듬기, 줄 바꿈, (릴리스 브랜치 모드) 비교 업데이트…
dgx-diagnose
nvidia
일반적인 DGX Station GB300 문제 진단 — CUDA 충돌, 잘못된 GPU 타겟팅, vLLM/SGLang 컨테이너 버그, MIG 상태 문제, NVLink/Fabric Manager 오류,…
aicr-managing-openvex
nvidia
Use when adding, updating, or removing CVE/GHSA suppressions in `.openvex.json` — the OpenVEX document consumed by the daily image vulnerability scan workflow.…
aicr-creating-slide-decks
nvidia
기술 개념이나 워크플로우에 대한 독립형 HTML 슬라이드 덱 또는 시각적 발표 자료(예: demos/*.html)를 만들 때 사용하세요. 전체 화면으로 표시하거나…
aicr-creating-guided-demos
nvidia
대화형 안내 데모 스크립트(demos/*.sh)를 라이브 또는 자기 주도 방식으로 Frame → Tell → Show → Close 패턴에 따라 구조화한다. "데모 스크립트", "안내…"와 같은 표현에 반응한다.
aicr-analyzing-snapshots
nvidia
AICR 스냅샷 YAML 파일을 분석하거나, 클러스터 상태를 검토하거나, 공급자 특성을 비교하거나, GPU/네트워크 토폴로지 인사이트를 추출할 때 사용합니다...