cuopt-debugging

par nvidia

Diagnostiquez les problèmes LP/MILP de cuOpt, y compris les erreurs, les résultats incorrects, les solutions irréalisables, les problèmes de performance et les codes de statut. À utiliser lorsque l'utilisateur dit…

npx skills add https://github.com/nvidia/cuopt-examples --skill cuopt-debugging

cuOpt Debugging Skill

Diagnose and fix issues with cuOpt LP/MILP solutions, errors, and performance.

Before You Start: Required Questions

Ask these to understand the problem:

  1. What's the symptom?

    • Error message?
    • Wrong/unexpected results?
    • Empty solution?
    • Performance too slow?
  2. What's the status?

    • problem.Status.name — what value does it show?
  3. Can you share?

    • The error message (exact text)
    • The code that produces it
    • Problem size (variables, constraints)

Quick Diagnosis by Symptom

"Solution is empty/None but status looks OK"

Most common cause: Wrong status string case

# ❌ WRONG - "OPTIMAL" never matches, silently fails
if problem.Status.name == "OPTIMAL":
    print(problem.ObjValue)  # Never runs!

# ✅ CORRECT - use PascalCase
if problem.Status.name in ["Optimal", "FeasibleFound"]:
    print(problem.ObjValue)

Diagnostic code:

print(f"Actual status: '{problem.Status.name}'")
print(f"Matches 'Optimal': {problem.Status.name == 'Optimal'}")
print(f"Matches 'OPTIMAL': {problem.Status.name == 'OPTIMAL'}")

"Objective value is wrong/zero"

Check if variables are actually used:

for var in problem.getVariables():
    print(f"{var.VariableName} = {var.Value}")
print(f"Objective: {problem.ObjValue}")

# Or with direct variable references
for var in [x, y, z]:
    print(f"{var.VariableName}: {var.getValue()}")

Common causes:

  • Constraints too restrictive (all zeros is feasible)
  • Objective coefficients have wrong sign
  • Wrong variable in objective

"Infeasible" status

For LP/MILP:

if problem.Status.name in ["PrimalInfeasible", "Infeasible"]:
    print("Problem has no feasible solution")
    # Review constraints for conflicts
    for c in problem.getConstraints():
        print(f"{c.ConstraintName}")

Common causes:

  • Conflicting constraints (x <= 5 AND x >= 10)
  • Bounds too tight
  • Missing a "slack" variable for soft constraints

"Integer variable has fractional value"

# Check how variable was defined
int_var = problem.addVariable(
    lb=0, ub=10,
    vtype=INTEGER,  # Must be INTEGER, not CONTINUOUS
    name="count"
)

# Also check if status is actually optimal
if problem.Status.name == "FeasibleFound":
    print("Warning: not fully optimal, may have fractional intermediate values")

"Unbounded" status

Problem has no finite optimum:

if problem.Status.name in ["DualInfeasible", "Unbounded"]:
    print("Problem is unbounded - objective can improve infinitely")

Common causes:

  • Missing variable upper/lower bounds
  • Constraint direction wrong (>= instead of <=)
  • Missing constraints

"Maximum recursion depth exceeded" when building expressions

Building large objectives or constraints with many chained + operations can hit Python recursion limits. Use LinearExpression instead:

from cuopt.linear_programming.problem import LinearExpression

# Instead of: expr = c1*v1 + c2*v2 + ... + cn*vn (many terms)
vars_list = [v1, v2, v3, ...]
coeffs_list = [c1, c2, c3, ...]
expr = LinearExpression(vars_list, coeffs_list, constant=0.0)
problem.setObjective(expr, sense=MINIMIZE)

See the LP/MILP "Building large expressions" section and reference models in the project for examples.

OutOfMemoryError

Check problem size:

print(f"Variables: {len(problem.getVariables())}")
print(f"Constraints: {len(problem.getConstraints())}")

Mitigations:

  • Reduce problem size
  • Use sparse constraint matrix
  • Set time limit to get partial solution

Status Code Reference

LP Status Values

StatusMeaning
OptimalFound optimal solution
PrimalFeasibleFound feasible but may not be optimal
PrimalInfeasibleNo feasible solution exists
DualInfeasibleProblem is unbounded
TimeLimitStopped due to time limit
IterationLimitStopped due to iteration limit
NumericalErrorNumerical issues encountered
NoTerminationSolver didn't converge

MILP Status Values

StatusMeaning
OptimalFound optimal solution
FeasibleFoundFound feasible, within gap tolerance
InfeasibleNo feasible solution exists
UnboundedProblem is unbounded
TimeLimitStopped due to time limit
NoTerminationNo solution found yet

Performance Debugging

Slow LP/MILP Solve

settings = SolverSettings()
settings.set_parameter("log_to_console", 1)  # See progress
settings.set_parameter("time_limit", 60)      # Don't wait forever

# For MILP, accept good-enough solution
settings.set_parameter("mip_relative_gap", 0.05)  # 5% gap

Check Solve Time

problem.solve(settings)
print(f"Solve time: {problem.SolveTime:.2f} seconds")

Diagnostic Checklist

□ Status checked with correct case (PascalCase)?
□ All variables have correct vtype (INTEGER vs CONTINUOUS)?
□ Constraint directions correct (<= vs >= vs ==)?
□ Objective sense correct (MINIMIZE vs MAXIMIZE)?
□ Variable bounds specified where needed?

Diagnostic Code Snippets

See resources/diagnostic_snippets.md for copy-paste diagnostic code:

  • Status checking
  • Variable inspection
  • Constraint analysis
  • Memory and performance checks

Interpreting Dual Values & Reduced Costs

When an LP/QP solve returns dual values and you need the decision read — which constraint is the binding bottleneck, what relaxing it is worth, and which unused option is the closest near-miss — see resources/interpreting_duals.md. (Integer models / MILP — and quadratic constraints — return no usable duals; that reference covers the fallback.)

When to Escalate

File a GitHub issue if:

  • Reproducible bug with minimal example
  • Include: cuOpt version, CUDA version, error message, minimal repro code

Plus de skills de nvidia

compileiq-debug
nvidia
Utilisez quand quelque chose ne va pas : Search() bloque, toutes les évaluations retournent INVALID_SCORE, les scores ne s'améliorent pas, chaque configuration retourne le même nombre, erreurs ptxas…
create-github-pr
nvidia
Créer des pull requests GitHub en utilisant l'interface en ligne de commande gh. Utiliser lorsque l'utilisateur souhaite créer une nouvelle PR, soumettre du code pour révision, ou ouvrir une pull request. Mots-clés de déclenchement -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Analyse les autres problèmes ouverts pour trouver ceux qu’une PR donnée pourrait également corriger ou casser accidentellement. Génère des opportunités de correctifs adjacents et des risques de contradiction avec fichier:ligne…
fhir-basics
nvidia
Apprend aux agents comment fonctionnent les API FHIR R4, quelles ressources sont disponibles, comment les interroger avec des paramètres de recherche, et comment analyser correctement tous les formats de réponse…
compileiq-validate-result
nvidia
Utiliser APRÈS qu'une recherche soit terminée et AVANT de réclamer un accélérateur ou d'expédier un ACF. Charge le CSV dump_results, extrait les K meilleurs candidats (mono-objectif)…
changelog-audit
nvidia
Auditer le CHANGELOG.md de Warp avant une publication : récupérer les entrées perdues, trier par impact utilisateur, affiner le langage des entrées, ajuster les retours à la ligne et (en mode branche de publication) mettre à jour la comparaison…
maintain-dynamic-plugins
nvidia
Maintenir les chargeurs de plugins dynamiques NeMo Relay, les manifestes, les SDK natifs Rust, le protocole worker gRPC, le SDK worker Python, la documentation, les tests et la couverture du workflow de publication
dgx-diagnose
nvidia
Diagnostiquer les problèmes courants du DGX Station GB300 — plantages CUDA, ciblage incorrect du GPU, bugs de conteneur vLLM/SGLang, problèmes d'état MIG, erreurs NVLink/Fabric Manager,…