cuopt-server-api-python

bởi nvidia

Máy chủ REST cuOpt — khởi động máy chủ, các điểm cuối, ví dụ về máy khách Python/curl. Sử dụng khi người dùng đang triển khai hoặc gọi API REST.

npx skills add https://github.com/nvidia/skills --skill cuopt-server-api-python

cuOpt Server — Deploy and client (Python/curl)

This skill covers starting the server and client examples (curl, Python). Server has no separate C API (clients can be any language).

Purpose

Use this skill when the user is deploying the cuOpt REST server or writing a client against it — choosing a deployment target, mapping a problem onto the HTTP endpoints, translating between Python-API and REST field names, or debugging a rejected payload.

Prerequisites

  • An NVIDIA GPU with a working CUDA driver (the server requires one; --gpus all for Docker).
  • cuopt-server installed, or Docker with the NVIDIA Container Toolkit. See the install skill.
  • Python clients need requests. No API key or auth token is required by the server itself.

Problem types supported

Problem typeSupported
Routing✓
LP✓
MILP✓
QP✗

Required questions

Ask these if not already clear:

  1. Problem type — Routing or LP/MILP? (QP not available via REST.)
  2. Deployment — Local, Docker, Kubernetes, or cloud?
  3. Client — Which language or tool will call the API (e.g. Python, curl, another service)?

Start server

# Development
python -m cuopt_server.cuopt_service --ip 0.0.0.0 --port 8000

# Docker — pick the tag matching your CUDA major version
docker run --gpus all -d -p 8000:8000 -e CUOPT_SERVER_PORT=8000 \
  nvidia/cuopt:latest-cu13

Use latest-cu12 or latest-cu13 to match your driver's CUDA major version (latest-cu13-ubi10 for a UBI10 base). Prefer these over the CUDA+Python-specific tags such as latest-cuda12.9-py3.13 — those track a single Python line and go stale when it stops receiving builds.

For production, pin rather than float: latest-* tags are mutable and can silently move to a different image. Use a full release tag (nvidia/cuopt:<release>-cuda<cuda>-py<python>) or an immutable digest (nvidia/cuopt@sha256:<digest>). Check the nvidia/cuopt registry for available tags.

Verify

Confirm the server is up by requesting GET /cuopt/health on the local port (e.g. http://localhost:8000/cuopt/health) — a healthy server returns HTTP 200.

Instructions

  1. POST to /cuopt/request → get reqId
  2. Poll /cuopt/solution/{reqId} until solution ready
  3. Parse response

Treat reqId as untrusted input: validate it (e.g. re.fullmatch(r"[A-Za-z0-9_-]{1,64}", req_id)) before interpolating it into the polling URL, and set an explicit timeout on every request.

Examples

import requests, time
SERVER = "http://localhost:8000"
HEADERS = {"Content-Type": "application/json", "CLIENT-VERSION": "custom"}
payload = {
    "cost_matrix_data": {"data": {"0": [[0,10,15],[10,0,12],[15,12,0]]}},
    "travel_time_matrix_data": {"data": {"0": [[0,10,15],[10,0,12],[15,12,0]]}},
    "task_data": {"task_locations": [1, 2], "demand": [[10, 20]], "task_time_windows": [[0,100],[0,100]], "service_times": [5, 5]},
    "fleet_data": {"vehicle_locations": [[0, 0]], "capacities": [[50]], "vehicle_time_windows": [[0, 200]]},
    "solver_config": {"time_limit": 5}
}
r = requests.post(f"{SERVER}/cuopt/request", json=payload, headers=HEADERS, timeout=30)
req_id = r.json()["reqId"]
# Poll: GET /cuopt/solution/{req_id}

Terminology: REST vs Python API

Python APIREST
order_locationstask_locations
set_order_time_windows()task_time_windows
service_timesservice_times

Use travel_time_matrix_data (not transit_time_matrix_data). Capacities: [[50, 50]] not [[50], [50]].

Troubleshooting

ErrorCauseSolution
422 Unprocessable EntityField name not in the schemaCheck names against the OpenAPI spec at /cuopt.yaml. Most common: transit_time_matrix_data → travel_time_matrix_data
422 on fleet_dataCapacities nested per vehicle instead of per dimensionUse [[50, 50]] (one inner list per capacity dimension), not [[50], [50]]
Connection refusedServer not up, or bound to a different interface/portcurl http://localhost:8000/cuopt/health; start with --ip 0.0.0.0 --port 8000
Docker container exits immediatelyNo GPU visible to the containerRun with --gpus all and confirm the NVIDIA Container Toolkit is installed
Polling never returns a solutionSolve exceeds the client's poll budgetRaise solver_config.time_limit and the poll loop count together

Capture the reqId and the full response body for any failed request — both are needed to diagnose server-side rejections.

Limitations

  • QP is not exposed over REST. Use the Python or C API for quadratic objectives.
  • The server ships no authentication or TLS. Anything that can reach the port can submit jobs. Put it behind a gateway and treat --server/base URLs as trusted-network endpoints only.
  • Solutions are retrieved by polling; there is no push/webhook delivery.
  • One request is solved at a time per server process; concurrency requires multiple replicas.

Runnable assets

Run from each asset directory (server must be running; scripts exit 0 if server unreachable). All use Python requests and accept --server (default http://localhost:8000):

See assets/README.md for overview.

Escalate

For contribution or build-from-source, see the developer skill.

Thêm skills từ nvidia

fhir-basics
nvidia
Dạy các tác nhân cách hoạt động của API FHIR R4, những tài nguyên có sẵn, cách truy vấn chúng với tham số tìm kiếm, và cách phân tích chính xác tất cả các định dạng phản hồi…
compileiq-validate-result
nvidia
Sử dụng SAU KHI tìm kiếm hoàn tất và TRƯỚC KHI yêu cầu tăng tốc hoặc gửi ACF. Tải tệp CSV dump_results, trích xuất các ứng viên top-K (đơn mục tiêu)…
changelog-audit
nvidia
Kiểm tra Warp CHANGELOG.md trước khi phát hành: khôi phục các mục bị mất, sắp xếp theo tác động người dùng, tinh chỉnh ngôn ngữ mục, xuống dòng và (chế độ nhánh phát hành) so sánh bump…
dgx-diagnose
nvidia
Chẩn đoán các sự cố thường gặp của DGX Station GB300 — lỗi CUDA, nhắm sai GPU, lỗi container vLLM/SGLang, vấn đề trạng thái MIG, lỗi NVLink/Fabric Manager,…
aicr-managing-openvex
nvidia
Use when adding, updating, or removing CVE/GHSA suppressions in `.openvex.json` — the OpenVEX document consumed by the daily image vulnerability scan workflow.…
aicr-creating-slide-decks
nvidia
Sử dụng khi xây dựng một bộ trình chiếu HTML độc lập hoặc điểm trình bày trực quan cho một khái niệm kỹ thuật hoặc quy trình làm việc (ví dụ: demos/*.html) — hiển thị toàn màn hình hoặc…
aicr-creating-guided-demos
nvidia
Tạo khung kịch bản demo có hướng dẫn tương tác (demos/*.sh), trực tiếp hoặc tự học, theo mẫu Frame → Tell → Show → Close. Kích hoạt khi có "demo script", "guided…
aicr-analyzing-snapshots
nvidia
Sử dụng khi phân tích tệp YAML snapshot AICR, xem xét trạng thái cụm, so sánh đặc điểm nhà cung cấp, trích xuất thông tin chi tiết về cấu trúc liên kết GPU/mạng, hoặc…