nemo-relay-debug-runtime-integration

от nvidia

Отладка проблем интеграции NeMo Relay на стороне приложения, таких как ошибки загрузки, неактивные области, отсутствующие события или проблемы с адаптивной/плагинной разводкой

npx skills add https://github.com/nvidia/nemo-relay --skill nemo-relay-debug-runtime-integration

Debug Runtime Integration

Use this skill when NeMo Relay is present in the application but something is not working. Start by proving which runtime layer is failing before changing configuration.

First Checks

  • Can the binding or native artifact load?
  • Is there an active scope when the failing call runs?
  • Is the work happening on the expected scope stack?
  • Is the subscriber/exporter/plugin configuration actually active?
  • Did the app choose the right public API layer: managed execute vs manual lifecycle vs typed wrappers vs adaptive/plugins?

Common Failure Classes

  • Python native extension missing
  • Go dynamic library not on the loader path
  • Node native addon not built or not loading
  • Execute call outside a scope
  • Missing events because registration never happened
  • Concurrency causing the wrong scope stack to be active
  • Adaptive component never initialized or config validation ignored

Embedded Troubleshooting Matrix

  • Rust build failure: run the narrowest core build first, then expand to the affected binding or workspace command.
  • Python import failure: rebuild the virtual environment and native extension with uv sync, then run a small Python test or import check from the same environment as the application.
  • Node.js addon failure: reinstall and rebuild the native addon from the Node binding package before debugging application code.
  • Go loader failure: build the release FFI shared library and point both the linker and runtime loader at the release output directory; use the macOS dynamic-library path variable (DYLD_LIBRARY_PATH) on macOS.
  • Scope stack empty: the work is outside an active scope or crossed a thread, task, goroutine, or worker boundary without the intended stack.
  • Work leaks across requests: separate requests are sharing one scope stack; create a fresh stack per independent request or agent.
  • Middleware missing or ordered incorrectly: check global vs scope-local registration, active scope ancestry, names, and priority values.
  • Subscriber missing events: register before the events are emitted; for scope-local subscribers, ensure the current scope is the owner or descendant.
  • Event fields missing: managed helpers populate semantic fields; manual lifecycle calls require explicit params for input, output, model names, and tool call IDs.
  • ATIF empty or mixed: register before work starts, use one exporter per run or clear between runs, and separate concurrent agents by root scope.
  • Provider payload conversion failure: convert non-JSON provider objects, SDK handles, callbacks, streams, or class instances with explicit codecs.
  • Plugin validation failure: validate config independently from runtime registration and check required fields, value types, defaults, and config source.
  • Adaptive behavior unchanged: confirm instrumentation emits events, the adaptive component is enabled, policy allows the behavior, and the call path reaches the configured component.
  • OpenTelemetry or OpenInference export failure: first identify the Relay version. For 0.6, check the separate exporter, http_binary versus grpc, endpoint, headers, target support, and an active Tokio runtime for native gRPC. For 0.7, check the typed projection, endpoint, headers, TLS, and target support; the subscriber owns the native gRPC runtime.
  • Callback succeeded but no lifecycle events appear: confirm the integration uses managed execute helpers or balanced manual start/end APIs, not only the underlying business callback.

Related Skills

  • nemo-relay-get-started
  • nemo-relay-instrument-context-isolation
  • nemo-relay-plugin-adaptive-tuning
  • nemo-relay-plugin-build

Больше skills от nvidia

compileiq-debug
nvidia
Используйте, когда что-то не так: Search() зависает, все оценки возвращают INVALID_SCORE, оценки не улучшаются, каждая конфигурация возвращает одно и то же число, ошибки ptxas…
create-github-pr
nvidia
Создание pull request'ов в GitHub с помощью gh CLI. Используйте, когда пользователь хочет создать новый PR, отправить код на ревью или открыть pull request. Ключевые слова для запуска —…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Сканирует другие открытые задачи, чтобы найти те, которые данный PR может исправить или случайно сломать. Выводит возможности смежных исправлений и риски противоречий с указанием файла:строки…
fhir-basics
nvidia
Обучает агентов работе с API FHIR R4, доступным ресурсам, запросам с параметрами поиска и корректному разбору всех форматов ответов…
compileiq-validate-result
nvidia
Используйте ПОСЛЕ завершения поиска и ДО применения ускорения или отправки ACF. Загружает CSV-файл dump_results, извлекает top-K кандидатов (однокритериальный)...
changelog-audit
nvidia
Аудит Warp CHANGELOG.md перед релизом: восстановление потерянных записей, сортировка по влиянию на пользователей, уточнение формулировок, перенос строк и (в режиме релизной ветки) обновление сравнения…
maintain-dynamic-plugins
nvidia
Поддержка загрузчиков динамических плагинов NeMo Relay, манифестов, нативных Rust SDK, протокола gRPC worker, Python worker SDK, документации, тестов и покрытия рабочего процесса релиза
dgx-diagnose
nvidia
Диагностика распространённых проблем DGX Station GB300 — сбои CUDA, ошибочное нацеливание на GPU, ошибки контейнеров vLLM/SGLang, проблемы состояния MIG, ошибки NVLink/Fabric Manager,…