vllm-pytorch-ci-triage

bởi pytorch

Phân loại một bản dựng CI vLLM Buildkite bị lỗi cho một PR nâng cấp phiên bản PyTorch, cô lập các hồi quy mới so với các lỗi đã tồn tại trên nhánh chính bằng cách so sánh với các bản dựng gần đây…

npx skills add https://github.com/pytorch/test-infra --skill vllm-pytorch-ci-triage

vLLM × PyTorch torch-nightly CI root-cause

Read-only root-cause analysis for the vLLM torch-nightly triage cron. The upstream triage job runs the A/B (torch nightly vs same-commit baseline), parses the logs, and hands off artifacts; this skill reads that artifact, root-causes each regressed cluster, groups by shared cause, routes each cause to the right repo, and writes its findings to findings.md / findings.json. It has no Buildkite/ClickHouse access, files no issues, and retries no jobs — tools are Read, Glob, Grep, Write.

Routing knowledge derived from the multi-week triage of vLLM PR #40077 (torch 2.12.0 + triton 3.7.0), which filed 16+ issues under umbrella pytorch/pytorch#180899.


Step 1: Read the cluster logs

The upstream triage job already produced the files — you do not run the parser or fetch anything. Your tools are Read, Glob, Grep, Write; there is no Python execution and no Buildkite / ClickHouse access. Read what is on disk in the triage input dir:

  • report.md — a concise, human-readable A/B summary and regressed clusters. Its Test-set regressions list deliberately renders each new failed test only as <test_id> — <pytest_exception_class>; it does not include tracebacks.
  • report.json — the structured companion: torch_nightly_build, baseline_build, commit, the regressed / both job lists, and regressed_tests. Each regressed_tests.new_failures entry is a complete FailedTest signature (test_id, pytest_exception_class, exception_chain, inline_message, test_is_infra). Each shared failure has complete torch_nightly and baseline FailedTest entries.
  • cluster-logs/<safe-cluster-name>.log — parsed pytest failures when available, bounded failure windows, plus an unconditional cleaned raw tail for each job-state regressed cluster.
  • both-cluster-logs/nightly_*.log — supplementary lossy raw nightly context for surfaced regressed_tests clusters. No baseline raw-context artifact is produced.

Cluster-log artifact format

Every file opens with a header:

# cluster: <cluster key>
# job: <job name>
# url: <buildkite job url>
# state: <state> exit_status: <n>

Job-state form — cluster-logs/<safe-cluster-name>.log uses capture_mode: nightly_failure_context and contains ranked failure windows followed by an unconditional cleaned ## raw_tail. When pytest parsing succeeds, a # parsed N failing test(s) section lists every failure's test ID, exception class, infra tag, and parsed exception chain before the windows. Window counts report candidates, emitted windows, and truncation. The header also includes job_is_infra: true|false; when it is true, treat the cluster as transient infrastructure rather than a torch regression. The windows are lossy evidence, not root-cause classifications or proof that a window belongs to a particular test.

Parsed pytest records for surfaced regressed_tests clusters, including complete nightly-only and shared nightly/baseline failures, are in report.json.regressed_tests.

both-cluster-logs/nightly_*.log uses capture_mode: both_failure_context and is supplementary lossy raw nightly context for the same surfaced clusters. No baseline_*.log artifact is produced.

Infra is not a torch regression. A job artifact tagged job_is_infra: true or a surfaced pytest failure tagged test_is_infra: true is transient infrastructure evidence; call it out as infra. Do not let a high-scoring infra message such as a low-GPU-memory startup failure override the explicit infra tag.

Step 2: NEW vs pre-existing

Classification is decided upstream by the torch-nightly vs same-commit-baseline A/B. Each job is already bucketed in the report:

  • regressed — fails on torch nightly, passes on the baseline → new, torch-attributable.
  • ordinary both entries — fail on both and remain pre-existing; ignore them.
  • regressed_tests — surfaced both clusters with a torch-nightly-only pytest failure; analyze them as new test-set regressions. Shared failures are context.
  • baseline_only — fails only on the baseline → ignore.

Rate new_failure_confidence (high/med/low) per group from its bucket plus the infra / agent-concentration evidence in the report. For scale definitions, see CONFIDENCE.md.

Step 3: Group remaining failures by root cause

Only genuinely new failures reach this step.

ONE group per root cause, not per job. From real data: 22 failing jobs grouped into 10 root causes.

For failed clusters, use the selected windows and raw tail. Scan upward from wrappers such as Engine core initialization failed. See root cause above. to find the real exception. For surfaced regressed_tests, use the complete nightly-only and shared records in report.json in combination with the both-cluster-logs/nightly_*.log.

Rate shared_root_cause_confidence (high/med/low) per member as you assign it to a group.

For grouping patterns, see GROUPING.md.

Step 4: Classify each group

Match each group's exception pattern against the routing cheat-sheet in ROUTING.md. Its Routing column is one of the three canonical values (pytorch/pytorch | vllm-project/vllm | infra) — the same set the triage workflow emits. Map routing to classification:

  • pytorch/pytorch → TORCH_REGRESSION
  • vllm-project/vllm → VLLM_REGRESSION
  • infra → not a regression — call it out as infra and do not file (see Step 1)

For scale definitions, see CONFIDENCE.md.

Gotchas the parser does NOT catch

The parser filters soft-fails, waiting_failed, never-ran jobs, marker/timestamp noise, and tags transient-infra signatures. The following are not filtered — apply them yourself when reading the cluster-logs/*.log files:

  • PyPI vs test channel: ERROR: No matching distribution found for torch==2.12.0 isn't infra — the release isn't on PyPI yet. It arrives as a non-pytest failure (treated as novel). Note it in the findings; it's not a bug to root-cause.
  • Python-only Installation job has multiple unrelated failure modes: (a) torch not on PyPI — expected, skip. (b) metadata is still not available after N attempts / precompiled wheel for commit X is available — vLLM's own precompiled-wheel infra hiccup, not torch. Both arrive as non-pytest failures (treated as novel) → ignore.
  • An infra-killed baseline job is not a baseline. The A/B buckets trust the baseline job's state. A baseline job hit by exit 125 / nvidia-container-cli (or otherwise never running the tests) still lands in BAD_STATES, so the same job failing on torch nightly is bucketed both (pre-existing) — masking a real regression rather than surfacing it. (The failing baseline was infra, not the same test.) When the baseline build has many B200 jobs killed by infra, do NOT trust a both verdict on those jobs; the pair is inconclusive because the baseline never ran the test. Flag it as inconclusive and recommend retrying the corresponding baseline job rather than concluding anything. The inverse mistake — treating a broken baseline as if the test passed there — produced a wrongful issue (#182549, retracted 2026-05-05).
  • Compile-on vs --enforce-eager CI gap: fake-kernel / Inductor stride bugs only surface when compile is on. Many gpt-oss CI lanes (tests/evals/gpt_oss/test_gpqa_correctness.py, --enforce-eager parametrizations) bypass torch.compile entirely and never trace the fake kernel. If a custom-op stride mismatch only shows up on the torch-bump test PR, the bug almost certainly exists on main too — vLLM CI is just hiding it. Call out this coverage gap in the findings.
  • Dockerfile.cpu seeds requirements/test/cpu.in from requirements/test/cuda.in (literal COPY ... cuda.in cpu.in), so the top-line --extra-index-url https://download.pytorch.org/whl/test/cu130 carries over to the CPU build. Combined with uv pip compile --torch-backend cpu (which forces stable cpu channel), torch 2.12 wheels go missing. Fix: sed-rewrite the index-url to whl/test/cpu AND drop --torch-backend cpu.
  • uv --torch-backend <name> overrides extra-index-url for torch. Only stable channels (cpu, cu128, etc.) are presets — there is no test-cpu preset. To pin torch to the test channel, use --extra-index-url explicitly (or UV_EXTRA_INDEX_URL env) and don't pass --torch-backend.

Thêm skills từ pytorch

zephyr
pytorch
Xây dựng và cấu hình ExecuTorch như một mô-đun Zephyr RTOS cho các bo mạch nhúng. Sử dụng khi thiết lập không gian làm việc Zephyr với ET, thêm hỗ trợ bo mạch (overlays,…
aoti-debug
pytorch
Gỡ lỗi các lỗi và sự cố của AOTInductor (AOTI). Sử dụng khi gặp lỗi segfault AOTI, lỗi không khớp thiết bị, lỗi tải hằng số, hoặc lỗi runtime từ…
skill-writer
pytorch
Hướng dẫn tạo Agent Skills có cấu trúc tốt cho Claude Code với các phương pháp hay nhất và xác thực. Bao gồm toàn bộ vòng đời Skill: phạm vi, cấu trúc tệp, xác thực YAML frontmatter, tổ chức nội dung và quy trình kiểm thử. Áp dụng quy tắc đặt tên nghiêm ngặt (chữ thường, dấu gạch ngang, tối đa 64 ký tự) và yêu cầu mô tả (kích hoạt cụ thể, loại tệp, mệnh đề "cái gì" và "khi nào"). Cung cấp mẫu cho các mẫu phổ biến bao gồm Skills chỉ đọc, Skills dựa trên tập lệnh và Skills nhiều tệp với...
triaging-issues
pytorch
Phân loại các vấn đề GitHub bằng cách chuyển đến các nhóm trực, áp dụng nhãn và đóng câu hỏi. Sử dụng khi xử lý các vấn đề PyTorch mới hoặc khi được yêu cầu phân loại một…
wheel-size-analyzer
pytorch
Phân tích kích thước wheel nightly của PyTorch trong một khoảng ngày bằng GitHub Actions artifacts API. Sử dụng khi theo dõi thay đổi kích thước binary, xác định kích thước wheel…
release-go-live-binary-build-matrix
pytorch
Cập nhật tools/scripts/generate_binary_build_matrix.py khi một bản phát hành PyTorch chính thức ra mắt. Nâng CURRENT_STABLE_VERSION lên phiên bản ổn định mới, thăng cấp…
pr-review
pytorch
Xem xét các pull request PyTorch về chất lượng mã, độ phủ kiểm thử, bảo mật và khả năng tương thích ngược. Sử dụng khi xem xét PR, khi được yêu cầu xem xét các thay đổi mã,…
qualcomm
pytorch
Xây dựng, kiểm thử hoặc phát triển backend QNN (Qualcomm AI Engine Direct). Sử dụng khi làm việc trên backends/qualcomm/, xây dựng QNN (sử dụng…