ci-metrics

작성자: pytorch

PyTorch CI, GitHub Actions, HUD, Grafana 및 인프라 메트릭을 조회합니다. 사용자가 CI 기간, 작업 실패, 대기 시간, 워크플로우 추세 등에 대해 질문할 때 사용하세요.

npx skills add https://github.com/pytorch/pytorch --skill ci-metrics

PyTorch CI Metrics

PyTorch CI and infrastructure metrics are exposed through Grafana. Use .claude/skills/ci-metrics/gcx-wrapper.sh for all Grafana access; it configures the PyTorch Grafana server, context, and authentication. Only users with write permission to the repo have access to Grafana. The authentication only provides read only access.

Requirements

The wrapper needs these tools on PATH:

  • gh - fetches the Grafana token and must be authenticated; if not, run gh auth login --hostname github.com --git-protocol ssh --web.
  • curl - downloads gcx and fetches the token from HUD.

On first use the wrapper downloads a pinned, checksum-verified gcx binary into a private cache (~/.cache/pytorch-ci-metrics/) and authenticates automatically. Nothing is installed on your PATH. If a tool is missing or gh is not authenticated, it exits with an error describing what to fix.

Datasources

Get the list of datasources available:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources list

The data contains metrics for many repos owned by the PyTorch repo. When possible, restrict queries to just the pytorch/pytorch repository.

CI and Test Run Data

CI and test run data are stored in grafana-clickhouse-datasource. List all the available tables:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources clickhouse list-tables

Important dataset:

  • GitHub webhook data
  • Tests
    • Database: tests
    • tests.all_test_runs - contains every test run. This is an extremely large table, so be considerate with filtering and timing.
    • Do not use tests.test_run_s3 as it contains partial data only.

To get additional guidance on common queries, clone https://github.com/pytorch/test-infra into a temporary directory and read the torchci folder.

Example Queries

Within the pytorch/pytorch repo on main, list the top most failing workflow jobs in the last 2 weeks:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources clickhouse query "
  SELECT name, count(DISTINCT id) AS failures
  FROM default.workflow_job
  WHERE conclusion = 'failure'
    AND completed_at >= now() - INTERVAL 2 WEEK
    AND repository_full_name = 'pytorch/pytorch'
    AND head_branch = 'main'
  GROUP BY name ORDER BY failures DESC LIMIT 10"

For a test file, how many times was it run in the last week? How many times did it pass or fail?

.claude/skills/ci-metrics/gcx-wrapper.sh datasources clickhouse query "
  SELECT
    file,
    classname,
    name,
    count() AS runs,
    countIf(failure_count = 0 AND error_count = 0 AND skipped_count = 0) AS successful,
    countIf(failure_count > 0 OR error_count > 0) AS fails,
    countIf(skipped_count > 0) AS skipped
  FROM tests.all_test_runs
  WHERE time_inserted >= now() - INTERVAL 7 DAY
    AND file = 'lazy/test_ts_opinfo.py'
  GROUP BY file, classname, name
  ORDER BY runs DESC"

CI Infrastructure

CI infrastructure metrics are stored in grafanacloud-pytorchci-prom. To get a better understanding of the data, clone these repositories in a temporary directory:

Example Queries

Which runner types have the deepest queue right now (jobs assigned but not yet running)?

.claude/skills/ci-metrics/gcx-wrapper.sh datasources prometheus query -d grafanacloud-prom 'topk(10, clamp_min(sum by (name) (gha_assigned_jobs) - sum by (name) (gha_running_jobs), 0))'

How many jobs were running per cluster over the last 6 hours, sampled every 30 minutes? Use --since/--step (or --from/--to) for a range query:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources prometheus query -d grafanacloud-prom 'sum by (cluster) (gha_running_jobs)' --since 6h --step 30m

pytorch의 다른 스킬

zephyr
pytorch
임베디드 보드용 Zephyr RTOS 모듈로 ExecuTorch를 빌드하고 구성합니다. ET로 Zephyr 워크스페이스를 설정하거나 보드 지원(오버레이 등)을 추가할 때 사용합니다.
aoti-debug
pytorch
AOTInductor(AOTI) 오류 및 충돌을 디버깅합니다. AOTI 세그폴트, 장치 불일치 오류, 상수 로딩 실패 또는 런타임 오류가 발생할 때 사용하세요.
skill-writer
pytorch
Claude Code를 위한 잘 구조화된 Agent Skill 생성 가이드로, 모범 사례와 검증을 포함합니다. Skill의 전체 수명 주기(범위 설정, 파일 구조, YAML 프론트매터 검증, 콘텐츠 구성, 테스트 절차)를 다룹니다. 엄격한 명명 규칙(소문자, 하이픈, 최대 64자)과 설명 요구 사항(특정 트리거, 파일 유형, "무엇" 및 "언제" 절)을 적용합니다. 읽기 전용 Skill, 스크립트 기반 Skill, 다중 파일 Skill 등 일반적인 패턴에 대한 템플릿을 제공합니다.
triaging-issues
pytorch
GitHub 이슈를 분류하여 온콜 팀에 라우팅하고, 레이블을 적용하며, 질문을 종료합니다. 새로운 PyTorch 이슈를 처리하거나 이슈 분류를 요청받았을 때 사용하세요.
wheel-size-analyzer
pytorch
PyTorch nightly wheel 크기를 GitHub Actions 아티팩트 API를 사용하여 날짜 범위에 걸쳐 분석합니다. 바이너리 크기 변경 추적, wheel 크기 식별에 사용합니다…
release-go-live-binary-build-matrix
pytorch
tools/scripts/generate_binary_build_matrix.py를 PyTorch 릴리스가 라이브될 때 업데이트합니다. CURRENT_STABLE_VERSION을 새로운 안정 버전으로 올리고, 해당…
pr-review
pytorch
PyTorch 풀 리퀘스트의 코드 품질, 테스트 커버리지, 보안 및 하위 호환성을 검토합니다. PR을 검토할 때, 코드 변경 사항을 검토하도록 요청받았을 때 사용합니다.
qualcomm
pytorch
QNN(Qualcomm AI Engine Direct) 백엔드를 빌드, 테스트 또는 개발합니다. backends/qualcomm/에서 작업하거나 QNN을 빌드할 때 사용합니다(계속…).