cuopt-server-api-python

작성자: nvidia

cuOpt REST 서버 — 서버 시작, 엔드포인트, Python/curl 클라이언트 예제. 사용자가 REST API를 배포하거나 호출할 때 사용합니다.

npx skills add https://github.com/nvidia/skills --skill cuopt-server-api-python

cuOpt Server — Deploy and client (Python/curl)

This skill covers starting the server and client examples (curl, Python). Server has no separate C API (clients can be any language).

Purpose

Use this skill when the user is deploying the cuOpt REST server or writing a client against it — choosing a deployment target, mapping a problem onto the HTTP endpoints, translating between Python-API and REST field names, or debugging a rejected payload.

Prerequisites

  • An NVIDIA GPU with a working CUDA driver (the server requires one; --gpus all for Docker).
  • cuopt-server installed, or Docker with the NVIDIA Container Toolkit. See the install skill.
  • Python clients need requests. No API key or auth token is required by the server itself.

Problem types supported

Problem typeSupported
Routing✓
LP✓
MILP✓
QP✗

Required questions

Ask these if not already clear:

  1. Problem type — Routing or LP/MILP? (QP not available via REST.)
  2. Deployment — Local, Docker, Kubernetes, or cloud?
  3. Client — Which language or tool will call the API (e.g. Python, curl, another service)?

Start server

# Development
python -m cuopt_server.cuopt_service --ip 0.0.0.0 --port 8000

# Docker — pick the tag matching your CUDA major version
docker run --gpus all -d -p 8000:8000 -e CUOPT_SERVER_PORT=8000 \
  nvidia/cuopt:latest-cu13

Use latest-cu12 or latest-cu13 to match your driver's CUDA major version (latest-cu13-ubi10 for a UBI10 base). Prefer these over the CUDA+Python-specific tags such as latest-cuda12.9-py3.13 — those track a single Python line and go stale when it stops receiving builds.

For production, pin rather than float: latest-* tags are mutable and can silently move to a different image. Use a full release tag (nvidia/cuopt:<release>-cuda<cuda>-py<python>) or an immutable digest (nvidia/cuopt@sha256:<digest>). Check the nvidia/cuopt registry for available tags.

Verify

Confirm the server is up by requesting GET /cuopt/health on the local port (e.g. http://localhost:8000/cuopt/health) — a healthy server returns HTTP 200.

Instructions

  1. POST to /cuopt/request → get reqId
  2. Poll /cuopt/solution/{reqId} until solution ready
  3. Parse response

Treat reqId as untrusted input: validate it (e.g. re.fullmatch(r"[A-Za-z0-9_-]{1,64}", req_id)) before interpolating it into the polling URL, and set an explicit timeout on every request.

Examples

import requests, time
SERVER = "http://localhost:8000"
HEADERS = {"Content-Type": "application/json", "CLIENT-VERSION": "custom"}
payload = {
    "cost_matrix_data": {"data": {"0": [[0,10,15],[10,0,12],[15,12,0]]}},
    "travel_time_matrix_data": {"data": {"0": [[0,10,15],[10,0,12],[15,12,0]]}},
    "task_data": {"task_locations": [1, 2], "demand": [[10, 20]], "task_time_windows": [[0,100],[0,100]], "service_times": [5, 5]},
    "fleet_data": {"vehicle_locations": [[0, 0]], "capacities": [[50]], "vehicle_time_windows": [[0, 200]]},
    "solver_config": {"time_limit": 5}
}
r = requests.post(f"{SERVER}/cuopt/request", json=payload, headers=HEADERS, timeout=30)
req_id = r.json()["reqId"]
# Poll: GET /cuopt/solution/{req_id}

Terminology: REST vs Python API

Python APIREST
order_locationstask_locations
set_order_time_windows()task_time_windows
service_timesservice_times

Use travel_time_matrix_data (not transit_time_matrix_data). Capacities: [[50, 50]] not [[50], [50]].

Troubleshooting

ErrorCauseSolution
422 Unprocessable EntityField name not in the schemaCheck names against the OpenAPI spec at /cuopt.yaml. Most common: transit_time_matrix_data → travel_time_matrix_data
422 on fleet_dataCapacities nested per vehicle instead of per dimensionUse [[50, 50]] (one inner list per capacity dimension), not [[50], [50]]
Connection refusedServer not up, or bound to a different interface/portcurl http://localhost:8000/cuopt/health; start with --ip 0.0.0.0 --port 8000
Docker container exits immediatelyNo GPU visible to the containerRun with --gpus all and confirm the NVIDIA Container Toolkit is installed
Polling never returns a solutionSolve exceeds the client's poll budgetRaise solver_config.time_limit and the poll loop count together

Capture the reqId and the full response body for any failed request — both are needed to diagnose server-side rejections.

Limitations

  • QP is not exposed over REST. Use the Python or C API for quadratic objectives.
  • The server ships no authentication or TLS. Anything that can reach the port can submit jobs. Put it behind a gateway and treat --server/base URLs as trusted-network endpoints only.
  • Solutions are retrieved by polling; there is no push/webhook delivery.
  • One request is solved at a time per server process; concurrency requires multiple replicas.

Runnable assets

Run from each asset directory (server must be running; scripts exit 0 if server unreachable). All use Python requests and accept --server (default http://localhost:8000):

See assets/README.md for overview.

Escalate

For contribution or build-from-source, see the developer skill.

nvidia의 다른 스킬

fhir-basics
nvidia
에이전트에게 FHIR R4 API의 작동 방식, 사용 가능한 리소스, 검색 매개변수를 사용한 쿼리 방법, 모든 응답 형식을 올바르게 파싱하는 방법을 가르칩니다…
compileiq-validate-result
nvidia
검색이 완료된 후, 속도 향상을 청구하거나 ACF를 발송하기 전에 사용합니다. dump_results CSV를 로드하고, 상위 K개 후보(단일 목표)를 추출합니다…
changelog-audit
nvidia
릴리스 전에 Warp CHANGELOG.md를 감사합니다: 누락된 항목 복구, 사용자 영향별 정렬, 항목 언어 다듬기, 줄 바꿈, (릴리스 브랜치 모드) 비교 업데이트…
dgx-diagnose
nvidia
일반적인 DGX Station GB300 문제 진단 — CUDA 충돌, 잘못된 GPU 타겟팅, vLLM/SGLang 컨테이너 버그, MIG 상태 문제, NVLink/Fabric Manager 오류,…
aicr-managing-openvex
nvidia
Use when adding, updating, or removing CVE/GHSA suppressions in `.openvex.json` — the OpenVEX document consumed by the daily image vulnerability scan workflow.…
aicr-creating-slide-decks
nvidia
기술 개념이나 워크플로우에 대한 독립형 HTML 슬라이드 덱 또는 시각적 발표 자료(예: demos/*.html)를 만들 때 사용하세요. 전체 화면으로 표시하거나…
aicr-creating-guided-demos
nvidia
대화형 안내 데모 스크립트(demos/*.sh)를 라이브 또는 자기 주도 방식으로 Frame → Tell → Show → Close 패턴에 따라 구조화한다. "데모 스크립트", "안내…"와 같은 표현에 반응한다.
aicr-analyzing-snapshots
nvidia
AICR 스냅샷 YAML 파일을 분석하거나, 클러스터 상태를 검토하거나, 공급자 특성을 비교하거나, GPU/네트워크 토폴로지 인사이트를 추출할 때 사용합니다...