debug-inference

bởi nvidia

Gỡ lỗi lý do thiết lập suy luận local hoặc bên ngoài bị lỗi. Sử dụng khi người dùng không thể kết nối đến máy chủ mô hình local, gặp vấn đề về URL cơ sở của nhà cung cấp, thấy…

npx skills add https://github.com/nvidia/openshell --skill debug-inference

Debug Inference

Diagnose inference as ordinary provider-authorized network traffic. OpenShell no longer supplies a managed inference route, rewrites request shapes, or selects a model. The application calls the provider's native endpoint and owns its base URL, model, request format, and timeout.

Use installed openshell --help output as the authority for command syntax. Refer to the published provider management guide and provider profile guide for current behavior.

Diagnostic Workflow

1. Confirm Gateway and Sandbox Context

openshell status
openshell gateway info
openshell sandbox get <sandbox>

For a host-local model server, host.openshell.internal identifies the machine running the gateway. It does not identify the operator's laptop when the gateway is remote. A server listening only on 127.0.0.1 may also be unreachable from a container; bind it to an address reachable from the gateway runtime.

2. Inspect the Provider and Its Profile

openshell provider get <provider>
openshell profile export <profile-id> -o yaml

Check that the profile:

  • Names the exact endpoint host, port, and protocol the client calls.
  • Allows the client binary.
  • Declares the credential key and intended authentication style.
  • Uses narrow HTTP rules when the provider should expose only part of an API.

For a custom or self-hosted OpenAI-compatible endpoint, import an endpoint-bearing profile. A base URL stored only in provider configuration does not authorize a new endpoint.

openshell profile lint -f ./provider-profile.yaml
openshell profile import -f ./provider-profile.yaml
openshell provider create --name <provider> --type <profile-id>

Add the required --credential KEY or --credential KEY=VALUE arguments shown by the profile. Never broaden endpoint policy merely to silence a credential binding error.

3. Confirm Attachment

openshell sandbox provider list <sandbox>
openshell sandbox provider attach <sandbox> <provider> --wait --timeout 30

Save the change's receipt_id and use openshell sandbox provider status <sandbox> <provider> --receipt <receipt-id> --wait --timeout 30 to check when it takes effect. Success confirms that the sandbox applied the credentials, policy, and environment for new processes. If the result is pending, failed, withheld, or superseded, inspect its reason before launching the client.

Launch the client after the attachment wait succeeds so it receives the updated environment:

openshell sandbox exec <sandbox> -- <client-command>

After updating an ordinary static provider, wait for that change and launch a new client. An existing process keeps its revision-scoped reference; readiness does not make the old reference resolve the replacement value. Diagnose managed-refresh credentials according to their own lifecycle.

Keep credentials and issued references out of diagnostic output. Acknowledged detach revokes future credential resolution and removes the reference from future process environments. Requests already forwarded may still finish:

openshell sandbox provider detach <sandbox> <provider> --wait --timeout 30

4. Verify Native Client Configuration

The application must use the real upstream contract:

  • Native provider base URL, not the retired managed virtual endpoint.
  • Real model ID, not a placeholder that OpenShell used to rewrite.
  • Native OpenAI, Anthropic, Vertex, or other provider request shape.
  • Application-owned timeout and retry settings.
  • The credential environment variable declared by the attached profile.

Probe the exact endpoint from a newly launched sandbox process. Start with a non-secret discovery endpoint when the provider offers one, then send a minimal inference request using the provider's documented API shape.

5. Interpret Common Failures

SymptomLikely causeFix
credential_placeholder_in_request_bodyA body reference is invalid/revoked, or classification metadata is unavailableCheck the controlled denial reason; remove the reference from conversation history or restore provider access. Do not enable body credential rewriting or bypass flags to send tool output. Unknown literals and valid issued placeholders pass unchanged, including the model provider’s own placeholder. Header resolution does not enable body rewriting.
A retired managed endpoint fails DNS resolutionClient still uses the removed managed endpointConfigure the provider's native base URL and attach an endpoint-bearing provider profile
Direct request is deniedMissing attachment, endpoint policy, HTTP rule, or binary authorizationInspect the attached provider profile and sandbox effective policy
credential_endpoint_mismatchCredential profile does not authorize the request recipientCorrect the host/port/path or import a narrowly scoped profile for the intended endpoint
request_authority_mismatchHTTP authority differs from the CONNECT destinationUse the same host and effective port in both authorities
Credential variable is absentProvider was not attached when this process launched, or profiles collide on a keyAttach the provider and launch a new process; resolve duplicate keys explicitly
Upstream rejects the model or bodyClient relied on removed model/request rewritingConfigure the real model and provider-native request format in the application
127.0.0.1 works on the host but not in the sandboxLoopback refers to different runtimeUse host.openshell.internal or another gateway-reachable endpoint and profile
Host-local request times outServer bind address, gateway topology, or host firewall blocks container-to-host trafficVerify the listener and permit only the required gateway network path and port

Host-Local Inference Checklist

For Ollama, LM Studio, vLLM, SGLang, TRT-LLM, and local NIM deployments:

  1. Verify the engine from the gateway host.
  2. Verify it listens on an address reachable from the gateway runtime.
  3. Import a custom profile naming host.openshell.internal and the actual port.
  4. Restrict the profile to the intended binaries and API paths.
  5. Create and attach the provider.
  6. Configure the application's base URL, model, and timeout.
  7. Probe the native endpoint from a newly launched sandbox process.

Reporting

Report:

  1. The active gateway and whether topology contributes to the failure.
  2. The provider, profile, attachment, endpoint, and client binary involved.
  3. The exact failed host, port, path, and request authority without secrets.
  4. Whether the client still relies on removed managed-routing behavior.
  5. The narrowest profile, attachment, or application configuration change that resolves the problem.

Thêm skills từ nvidia

fhir-basics
nvidia
Dạy các tác nhân cách hoạt động của API FHIR R4, những tài nguyên có sẵn, cách truy vấn chúng với tham số tìm kiếm, và cách phân tích chính xác tất cả các định dạng phản hồi…
compileiq-validate-result
nvidia
Sử dụng SAU KHI tìm kiếm hoàn tất và TRƯỚC KHI yêu cầu tăng tốc hoặc gửi ACF. Tải tệp CSV dump_results, trích xuất các ứng viên top-K (đơn mục tiêu)…
changelog-audit
nvidia
Kiểm tra Warp CHANGELOG.md trước khi phát hành: khôi phục các mục bị mất, sắp xếp theo tác động người dùng, tinh chỉnh ngôn ngữ mục, xuống dòng và (chế độ nhánh phát hành) so sánh bump…
dgx-diagnose
nvidia
Chẩn đoán các sự cố thường gặp của DGX Station GB300 — lỗi CUDA, nhắm sai GPU, lỗi container vLLM/SGLang, vấn đề trạng thái MIG, lỗi NVLink/Fabric Manager,…
aicr-managing-openvex
nvidia
Use when adding, updating, or removing CVE/GHSA suppressions in `.openvex.json` — the OpenVEX document consumed by the daily image vulnerability scan workflow.…
aicr-creating-slide-decks
nvidia
Sử dụng khi xây dựng một bộ trình chiếu HTML độc lập hoặc điểm trình bày trực quan cho một khái niệm kỹ thuật hoặc quy trình làm việc (ví dụ: demos/*.html) — hiển thị toàn màn hình hoặc…
aicr-creating-guided-demos
nvidia
Tạo khung kịch bản demo có hướng dẫn tương tác (demos/*.sh), trực tiếp hoặc tự học, theo mẫu Frame → Tell → Show → Close. Kích hoạt khi có "demo script", "guided…
aicr-analyzing-snapshots
nvidia
Sử dụng khi phân tích tệp YAML snapshot AICR, xem xét trạng thái cụm, so sánh đặc điểm nhà cung cấp, trích xuất thông tin chi tiết về cấu trúc liên kết GPU/mạng, hoặc…