compileiq-bootstrap

par nvidia

À utiliser lors du démarrage d'un nouveau projet CompileIQ, en cas de timeout de socket, ou avant d'exécuter toute autre compétence compileiq-*. Vérifie CUDA 13.3+, ptxas, l'accès GPU,…

npx skills add https://github.com/nvidia/compileiq --skill compileiq-bootstrap

compileiq-bootstrap

Validate that a host can run CompileIQ. No installation of system libraries (that's distro-specific and out of scope). No state files. Just a checklist that either passes or prints the precise next fix.

When

  • First-time setup on a new host or container.
  • Search().start() hangs on the first evaluation.
  • After upgrading CUDA, the compileiq wheel, or the Python venv.

Steps

1. CUDA toolkit 13.3+

CompileIQ's --apply-controls mechanism requires CUDA 13.3 or later (see examples/compilers/nvbench_example/optimize_reduction.py:302).

nvcc --version | grep -E "release (1[3-9]|[2-9][0-9])\.[3-9]"
ptxas --version | grep -E "V(1[3-9]|[2-9][0-9])\.[3-9]"

If either command exits non-zero, install or upgrade the CUDA toolkit and put /usr/local/cuda/bin on PATH and /usr/local/cuda/lib64 on LD_LIBRARY_PATH.

2. GPU is visible

nvidia-smi --query-gpu=name,compute_cap --format=csv

CompileIQ targets compute capability 9.0 (Hopper / H100) and 10.0 (Blackwell / B200) most aggressively, but works on any GPU PTXAS supports for the chosen arch.

3. CompileIQ imports resolve

One shot that covers everything callers will need:

python -c "
from compileiq.ciq import Search
from compileiq.types import INVALID_SCORE, BASELINE_CONFIG, WorkerTypes, ProblemType, SearchConfiguration
from compileiq.search_spaces.compilers import PtxasSearchSpace, NvccSearchSpace, LocalSearchSpaceBin
from compileiq.utils.helpers import save_compiler_config, load_compiler_config
from compileiq.worker import MultiProcessWorker, IsoMultiProcessWorker, RayWorker, AsyncWorker
print('imports OK')
"

If this fails, run pip install compileiq (or pip install -e . from a source checkout) and re-run.

4. Search-space resolution round-trip

This catches network or air-gapped issues that otherwise surface as socket-timeout hangs deep inside the first evaluation:

python -c "
from compileiq.search_spaces.compilers import PtxasSearchSpace
p = PtxasSearchSpace().retrieve()
assert p.exists() and p.stat().st_size > 0, p
print(f'resolved: {p}')
"

The first run downloads from GitHub releases and caches under ~/.cache/compileiq/<tag>/. Subsequent runs hit the cache.

5. Env vars to know

VariableDefaultWhat it controls
CIQ_SOCKET_TIMEOUT20Seconds to wait on IPC with the core. Raise to 60-120 for large search spaces; raise much higher if you regularly see a hang on the first evaluation.
CIQ_KEEP_CACHEunset (0)Set to 1/true/yes to keep .cache files after a run for post-mortem replay.
CIQ_PROCESS_MODEforkserverforkserver, fork, or spawn. Switch to spawn if forkserver fails on a constrained host. IsoMultiProcessWorker defaults to fork regardless.
CIQ_SEARCH_SPACES_DIRunsetPath to a local mirror containing manifest.json plus the referenced .bin files. Set this on air-gapped hosts to skip network.
CIQ_SEARCH_SPACES_REPONVIDIA/CompileIQOverride the GitHub repo that release-backed search-space resolution queries. Useful for staging or forks.
CIQ_SS_TAG_PREFIXsearch-spaces-Tag prefix the resolver uses when tag="latest". Rarely needs changing.

Self-test

bash scripts/check_env.sh

Exits 0 if every step passes. Exits with a non-zero count of failures and prints, for each failure, the precise command the user should run next.

Gotchas

  • Socket timeout on first eval is not BLAS. Current shipped binaries do not require BLAS/LAPACK. If a hang occurs:
    1. Raise CIQ_SOCKET_TIMEOUT=120 and retry.
    2. If still hangs, mirror search spaces locally and set CIQ_SEARCH_SPACES_DIR=/path/to/mirror.
    3. If still hangs, switch process mode: CIQ_PROCESS_MODE=spawn.
  • Don't install system libraries in this skill. Distro-specific package installs require sudo and break in containers and on shared clusters. If nvcc, ptxas, or a Python interpreter is missing, surface the error and let the user decide how to install.

Next

  • For attention workloads: compileiq-search-space (variant="att" by default).
  • Before paying for a full search: try compileiq-booster-pack.
  • Writing an objective function: compileiq-author-objective.

Plus de skills de nvidia

compileiq-debug
nvidia
Utilisez quand quelque chose ne va pas : Search() bloque, toutes les évaluations retournent INVALID_SCORE, les scores ne s'améliorent pas, chaque configuration retourne le même nombre, erreurs ptxas…
create-github-pr
nvidia
Créer des pull requests GitHub en utilisant l'interface en ligne de commande gh. Utiliser lorsque l'utilisateur souhaite créer une nouvelle PR, soumettre du code pour révision, ou ouvrir une pull request. Mots-clés de déclenchement -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Analyse les autres problèmes ouverts pour trouver ceux qu’une PR donnée pourrait également corriger ou casser accidentellement. Génère des opportunités de correctifs adjacents et des risques de contradiction avec fichier:ligne…
fhir-basics
nvidia
Apprend aux agents comment fonctionnent les API FHIR R4, quelles ressources sont disponibles, comment les interroger avec des paramètres de recherche, et comment analyser correctement tous les formats de réponse…
compileiq-validate-result
nvidia
Utiliser APRÈS qu'une recherche soit terminée et AVANT de réclamer un accélérateur ou d'expédier un ACF. Charge le CSV dump_results, extrait les K meilleurs candidats (mono-objectif)…
changelog-audit
nvidia
Auditer le CHANGELOG.md de Warp avant une publication : récupérer les entrées perdues, trier par impact utilisateur, affiner le langage des entrées, ajuster les retours à la ligne et (en mode branche de publication) mettre à jour la comparaison…
maintain-dynamic-plugins
nvidia
Maintenir les chargeurs de plugins dynamiques NeMo Relay, les manifestes, les SDK natifs Rust, le protocole worker gRPC, le SDK worker Python, la documentation, les tests et la couverture du workflow de publication
dgx-diagnose
nvidia
Diagnostiquer les problèmes courants du DGX Station GB300 — plantages CUDA, ciblage incorrect du GPU, bugs de conteneur vLLM/SGLang, problèmes d'état MIG, erreurs NVLink/Fabric Manager,…