distributed-triage

작성자: pytorch

온콜:분산 큐의 이슈를 하위 분류하여 분산 모듈 레이블을 할당하고, 하위 온콜로 라우팅하며, 분류 완료로 표시합니다. 이슈가 다음과 같은 경우 사용합니다…

npx skills add https://github.com/pytorch/pytorch --skill distributed-triage

Distributed Issue Triage Sub-Skill

This sub-skill picks up where the PT-level triage bot leaves off. It processes issues that already have the oncall: distributed label and performs second-level triage: routing to a distributed sub-oncall, classifying by module, and marking triaged.

Contents

Distributed labels reference: See distributed-labels.json for the labels this skill is allowed to apply. ONLY apply labels from this file.

Distributed triage rubric: See distributed-rubric.md for detailed routing guidance, module classification signals, and confidence calibration.

Response templates: See templates.json for distributed-specific comment templates.


MCP Tools Available

Use these GitHub MCP tools for triage:

ToolPurpose
mcp__github__get_issueGet issue details and existing labels
mcp__github__get_issue_commentsGet existing issue comments
mcp__github__update_issueApply labels or close issues
mcp__github__add_issue_commentAdd comment (only for reproduction requests or mislabel flags)
mcp__github__search_issuesFind similar issues for context

Comment Deduplication

Before adding any issue comment:

  1. Read the existing comments with mcp__github__get_issue_comments.
  2. Check whether the triage bot has already posted the same template or a substantially equivalent request/explanation.
  3. If a duplicate exists, do not add another comment. Continue with any non-comment actions that are still needed, such as labels.

Treat a comment as duplicate even if the wording differs slightly or an older template version was used. For distributed triage, this includes an existing distributed reproduction request or an existing "not distributed" notice.


Distributed Triage Steps

0) Already Triaged by Human?

A human has fully classified the issue only when it has BOTH:

  1. Any module: label listed in distributed-labels.json, AND
  2. One of the sub-oncall labels: oncall: distributed parallelisms, oncall: distributed infra, or oncall: distributed checkpointing.

If both are present:

  • Add bot-triaged + triaged labels (the human classification is complete and confident)
  • STOP — a human already classified this issue.

If only one is present (a module label without a sub-oncall, or a sub-oncall without a module label), triage is incomplete — proceed to Step 1. The PT-level triage bot can apply distributed module labels alongside oncall: distributed, but it does not pick the sub-oncall; that is your job.

This step alone should clear a large portion of the backlog.

1) Is This Actually a Distributed Issue?

Read the issue title, description, and comments. Determine whether the issue is actually related to distributed training.

Signs it is NOT a distributed issue:

  • Single-GPU issue with no distributed code (e.g., torch.nn on one GPU, CUDA OOM on one device)
  • Build/packaging issue (e.g., undefined symbol: ncclAlltoAll at import torch with no distributed code)
  • Pure torch.compile issue with no distributed component
  • Issue about a domain library (vision, text, audio) that happens to mention "distributed"

If NOT a distributed issue:

  1. Add triage review + bot-triaged labels
  2. Post a comment using the not_distributed template from templates.json, unless an equivalent "not distributed" comment already exists
  3. Do NOT remove oncall: distributed — let the human oncall re-route
  4. STOP

2) Route to Distributed Sub-Oncall

Each issue carries exactly ONE sub-oncall label. If the issue already has one of the three sub-oncall labels (oncall: distributed parallelisms, oncall: distributed infra, or oncall: distributed checkpointing), keep it as-is — do NOT add a second sub-oncall, even if your own classification would have picked a different one. Use the existing sub-oncall to decide the next step (continue to Step 3 if it's oncall: distributed parallelisms; otherwise add bot-triaged and STOP per the rules below).

If no sub-oncall is present, apply exactly one based on the routing rules in distributed-rubric.md:

Sub-Oncall LabelWhen to Apply
oncall: distributed parallelismsFSDP, DDP, DTensor, tensor parallel, context parallel, pipeline parallel. This is the default when unsure.
oncall: distributed infrac10d, process groups, collectives, NCCL/Gloo/MPI backends, elastic/torchrun, RPC, stores, distributed tools, DeviceMesh, symmetric memory
oncall: distributed checkpointingDistributed checkpoint save/load, DCP, state_dict utilities, async checkpointing

Use the routing decision tree and edge cases in distributed-rubric.md Section 1 to determine the correct sub-oncall.

After routing to oncall: distributed infra or oncall: distributed checkpointing:

  • Add bot-triaged (the sub-oncall routing is a confident, complete outcome)
  • STOP — the sub-oncall team owns further triage

After routing to oncall: distributed parallelisms:

  • Continue to Step 3 for module classification

3) Classify Module

From the issue description, comments, code snippets, and stack traces, classify into one or more distributed modules. Consult the module classification signals in distributed-rubric.md.

Confidence-based actions:

ConfidenceCriteriaAction
HIGH or MEDIUMExplicit module mention, obvious API usage, or probable module based on contextAdd module: label(s) + bot-triaged + triaged
LOWCannot determine module — vague description, no code, no stack traceAdd triage review + bot-triaged (no triaged — punting to a human)

Rules:

  • You can apply multiple module labels when the issue spans modules (e.g., module: fsdp + module: dtensor for FSDP2 issues that hit DTensor bugs).
  • When an issue has oncall: pt2 already applied, do NOT remove it. Add distributed module labels alongside it.
  • When the module is unclear, add triage review + bot-triaged — do NOT guess a module label.

4) Type Labels

If the issue is not a bug report, add the appropriate type label:

  • feature — wholly new functionality that does not exist today in any form
  • enhancement — improvement to something that already works (e.g., performance optimization, better error messages, adding a native backend for an op that already runs via fallback)

Most distributed issues are bug reports — do not add a type label for bugs. If the issue says the operation "currently works" or "falls back to" a slower path, that is enhancement, not feature. If the enhancement is about performance, also add module: performance.

5) High Priority — REQUIRES HUMAN REVIEW

CRITICAL: If you believe an issue is high priority, you MUST:

  1. Add triage review label and do NOT add bot-triaged

Do NOT directly add high priority without human confirmation.

High priority criteria for distributed issues:

  • Crash / segfault / illegal memory access in distributed code
  • Silent correctness issue (wrong results from collectives, incorrect gradient sync)
  • Regression from a prior version (e.g., FSDP worked in 2.x, broken in 2.y)
  • Hang affecting multi-node training (NCCL timeout, deadlock in collectives)
  • Data corruption during distributed checkpointing
  • Internal assert failure in c10d or process group code
  • Many users affected or core distributed component impacted

6) Missing Reproduction

If the issue lacks a minimal reproduction script:

  1. Add needs reproduction + bot-triaged labels
  2. Post a comment using the needs_distributed_reproduction template from templates.json, unless an equivalent distributed reproduction request already exists

Do NOT request reproduction when:

  • The issue already has a code snippet, script, or steps that someone could follow to reproduce
  • The issue is a feature request (no repro needed)
  • A multi-node script is provided (that counts as reproduction even if you can't run it locally)

Constraints

DO NOT:

  • Close issues (only the PT-level bot or humans close issues)
  • Remove existing labels — only add labels
  • Remove oncall: distributed — it stays even if the issue is mislabeled
  • Remove oncall: pt2 — if already present, keep it
  • Remove bot-triaged or triaged — they are applied by the parent skill and must stay
  • Add triaged when you are NOT confident in the classification — i.e. any time the action also applies triage review or needs reproduction, or in the §5 high-priority flow
  • Add labels not in distributed-labels.json
  • Add comments to issues except when using the templates in Step 1 (mislabel) or Step 6 (reproduction)
  • Add a comment when the bot has already posted the same template or a substantially equivalent message on the issue
  • Assign issues to users
  • Add high priority directly — use triage review and let humans decide

DO:

  • Be conservative — when in doubt, add triage review for human attention
  • Add bot-triaged whenever the bot has processed the issue, regardless of confidence. Pair with triage review for LOW-confidence or uncertain cases so the cron sweep won't re-pick it. (Exception: §5 high-priority flow intentionally omits bot-triaged.)
  • Add triaged ONLY when you reach a confident, complete classification: a human already classified it (Step 0), a confident sub-oncall routing (Step 2), or a HIGH/MEDIUM-confidence module classification (Step 3).
  • Always add a sub-oncall label (Step 2) before module labels (Step 3)
  • Read the full issue including comments before classifying
  • Read existing comments before every comment action and skip duplicate bot messages
  • Check the rubric's "Common Mislabel Traps" section before finalizing

pytorch의 다른 스킬

zephyr
pytorch
임베디드 보드용 Zephyr RTOS 모듈로 ExecuTorch를 빌드하고 구성합니다. ET로 Zephyr 워크스페이스를 설정하거나 보드 지원(오버레이 등)을 추가할 때 사용합니다.
aoti-debug
pytorch
AOTInductor(AOTI) 오류 및 충돌을 디버깅합니다. AOTI 세그폴트, 장치 불일치 오류, 상수 로딩 실패 또는 런타임 오류가 발생할 때 사용하세요.
skill-writer
pytorch
Claude Code를 위한 잘 구조화된 Agent Skill 생성 가이드로, 모범 사례와 검증을 포함합니다. Skill의 전체 수명 주기(범위 설정, 파일 구조, YAML 프론트매터 검증, 콘텐츠 구성, 테스트 절차)를 다룹니다. 엄격한 명명 규칙(소문자, 하이픈, 최대 64자)과 설명 요구 사항(특정 트리거, 파일 유형, "무엇" 및 "언제" 절)을 적용합니다. 읽기 전용 Skill, 스크립트 기반 Skill, 다중 파일 Skill 등 일반적인 패턴에 대한 템플릿을 제공합니다.
triaging-issues
pytorch
GitHub 이슈를 분류하여 온콜 팀에 라우팅하고, 레이블을 적용하며, 질문을 종료합니다. 새로운 PyTorch 이슈를 처리하거나 이슈 분류를 요청받았을 때 사용하세요.
wheel-size-analyzer
pytorch
PyTorch nightly wheel 크기를 GitHub Actions 아티팩트 API를 사용하여 날짜 범위에 걸쳐 분석합니다. 바이너리 크기 변경 추적, wheel 크기 식별에 사용합니다…
release-go-live-binary-build-matrix
pytorch
tools/scripts/generate_binary_build_matrix.py를 PyTorch 릴리스가 라이브될 때 업데이트합니다. CURRENT_STABLE_VERSION을 새로운 안정 버전으로 올리고, 해당…
pr-review
pytorch
PyTorch 풀 리퀘스트의 코드 품질, 테스트 커버리지, 보안 및 하위 호환성을 검토합니다. PR을 검토할 때, 코드 변경 사항을 검토하도록 요청받았을 때 사용합니다.
qualcomm
pytorch
QNN(Qualcomm AI Engine Direct) 백엔드를 빌드, 테스트 또는 개발합니다. backends/qualcomm/에서 작업하거나 QNN을 빌드할 때 사용합니다(계속…).