nemo-relay-debug-runtime-integration

par nvidia

Déboguer les problèmes d'intégration NeMo Relay côté application, tels que les échecs de chargement, les scopes inactifs, les événements manquants ou les problèmes de câblage adaptatif/plugin.

npx skills add https://github.com/nvidia/nemo-relay --skill nemo-relay-debug-runtime-integration

Debug Runtime Integration

Use this skill when NeMo Relay is present in the application but something is not working. Start by proving which runtime layer is failing before changing configuration.

First Checks

  • Can the binding or native artifact load?
  • Is there an active scope when the failing call runs?
  • Is the work happening on the expected scope stack?
  • Is the subscriber/exporter/plugin configuration actually active?
  • Did the app choose the right public API layer: managed execute vs manual lifecycle vs typed wrappers vs adaptive/plugins?

Common Failure Classes

  • Python native extension missing
  • Go dynamic library not on the loader path
  • Node native addon not built or not loading
  • Execute call outside a scope
  • Missing events because registration never happened
  • Concurrency causing the wrong scope stack to be active
  • Adaptive component never initialized or config validation ignored

Embedded Troubleshooting Matrix

  • Rust build failure: run the narrowest core build first, then expand to the affected binding or workspace command.
  • Python import failure: rebuild the virtual environment and native extension with uv sync, then run a small Python test or import check from the same environment as the application.
  • Node.js addon failure: reinstall and rebuild the native addon from the Node binding package before debugging application code.
  • Go loader failure: build the release FFI shared library and point both the linker and runtime loader at the release output directory; use the macOS dynamic-library path variable (DYLD_LIBRARY_PATH) on macOS.
  • Scope stack empty: the work is outside an active scope or crossed a thread, task, goroutine, or worker boundary without the intended stack.
  • Work leaks across requests: separate requests are sharing one scope stack; create a fresh stack per independent request or agent.
  • Middleware missing or ordered incorrectly: check global vs scope-local registration, active scope ancestry, names, and priority values.
  • Subscriber missing events: register before the events are emitted; for scope-local subscribers, ensure the current scope is the owner or descendant.
  • Event fields missing: managed helpers populate semantic fields; manual lifecycle calls require explicit params for input, output, model names, and tool call IDs.
  • ATIF empty or mixed: register before work starts, use one exporter per run or clear between runs, and separate concurrent agents by root scope.
  • Provider payload conversion failure: convert non-JSON provider objects, SDK handles, callbacks, streams, or class instances with explicit codecs.
  • Plugin validation failure: validate config independently from runtime registration and check required fields, value types, defaults, and config source.
  • Adaptive behavior unchanged: confirm instrumentation emits events, the adaptive component is enabled, policy allows the behavior, and the call path reaches the configured component.
  • OpenTelemetry or OpenInference export failure: first identify the Relay version. For 0.6, check the separate exporter, http_binary versus grpc, endpoint, headers, target support, and an active Tokio runtime for native gRPC. For 0.7, check the typed projection, endpoint, headers, TLS, and target support; the subscriber owns the native gRPC runtime.
  • Callback succeeded but no lifecycle events appear: confirm the integration uses managed execute helpers or balanced manual start/end APIs, not only the underlying business callback.

Related Skills

  • nemo-relay-get-started
  • nemo-relay-instrument-context-isolation
  • nemo-relay-plugin-adaptive-tuning
  • nemo-relay-plugin-build

Plus de skills de nvidia

compileiq-debug
nvidia
Utilisez quand quelque chose ne va pas : Search() bloque, toutes les évaluations retournent INVALID_SCORE, les scores ne s'améliorent pas, chaque configuration retourne le même nombre, erreurs ptxas…
create-github-pr
nvidia
Créer des pull requests GitHub en utilisant l'interface en ligne de commande gh. Utiliser lorsque l'utilisateur souhaite créer une nouvelle PR, soumettre du code pour révision, ou ouvrir une pull request. Mots-clés de déclenchement -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Analyse les autres problèmes ouverts pour trouver ceux qu’une PR donnée pourrait également corriger ou casser accidentellement. Génère des opportunités de correctifs adjacents et des risques de contradiction avec fichier:ligne…
fhir-basics
nvidia
Apprend aux agents comment fonctionnent les API FHIR R4, quelles ressources sont disponibles, comment les interroger avec des paramètres de recherche, et comment analyser correctement tous les formats de réponse…
compileiq-validate-result
nvidia
Utiliser APRÈS qu'une recherche soit terminée et AVANT de réclamer un accélérateur ou d'expédier un ACF. Charge le CSV dump_results, extrait les K meilleurs candidats (mono-objectif)…
changelog-audit
nvidia
Auditer le CHANGELOG.md de Warp avant une publication : récupérer les entrées perdues, trier par impact utilisateur, affiner le langage des entrées, ajuster les retours à la ligne et (en mode branche de publication) mettre à jour la comparaison…
maintain-dynamic-plugins
nvidia
Maintenir les chargeurs de plugins dynamiques NeMo Relay, les manifestes, les SDK natifs Rust, le protocole worker gRPC, le SDK worker Python, la documentation, les tests et la couverture du workflow de publication
dgx-diagnose
nvidia
Diagnostiquer les problèmes courants du DGX Station GB300 — plantages CUDA, ciblage incorrect du GPU, bugs de conteneur vLLM/SGLang, problèmes d'état MIG, erreurs NVLink/Fabric Manager,…