ci-metrics

作者: pytorch

查詢 PyTorch CI、GitHub Actions、HUD、Grafana 以及基礎設施指標。當使用者詢問關於 CI 持續時間、任務失敗、佇列時間、工作流程趨勢等問題時使用。

npx skills add https://github.com/pytorch/pytorch --skill ci-metrics

PyTorch CI Metrics

PyTorch CI and infrastructure metrics are exposed through Grafana. Use .claude/skills/ci-metrics/gcx-wrapper.sh for all Grafana access; it configures the PyTorch Grafana server, context, and authentication. Only users with write permission to the repo have access to Grafana. The authentication only provides read only access.

Requirements

The wrapper needs these tools on PATH:

  • gh - fetches the Grafana token and must be authenticated; if not, run gh auth login --hostname github.com --git-protocol ssh --web.
  • curl - downloads gcx and fetches the token from HUD.

On first use the wrapper downloads a pinned, checksum-verified gcx binary into a private cache (~/.cache/pytorch-ci-metrics/) and authenticates automatically. Nothing is installed on your PATH. If a tool is missing or gh is not authenticated, it exits with an error describing what to fix.

Datasources

Get the list of datasources available:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources list

The data contains metrics for many repos owned by the PyTorch repo. When possible, restrict queries to just the pytorch/pytorch repository.

CI and Test Run Data

CI and test run data are stored in grafana-clickhouse-datasource. List all the available tables:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources clickhouse list-tables

Important dataset:

  • GitHub webhook data
  • Tests
    • Database: tests
    • tests.all_test_runs - contains every test run. This is an extremely large table, so be considerate with filtering and timing.
    • Do not use tests.test_run_s3 as it contains partial data only.

To get additional guidance on common queries, clone https://github.com/pytorch/test-infra into a temporary directory and read the torchci folder.

Example Queries

Within the pytorch/pytorch repo on main, list the top most failing workflow jobs in the last 2 weeks:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources clickhouse query "
  SELECT name, count(DISTINCT id) AS failures
  FROM default.workflow_job
  WHERE conclusion = 'failure'
    AND completed_at >= now() - INTERVAL 2 WEEK
    AND repository_full_name = 'pytorch/pytorch'
    AND head_branch = 'main'
  GROUP BY name ORDER BY failures DESC LIMIT 10"

For a test file, how many times was it run in the last week? How many times did it pass or fail?

.claude/skills/ci-metrics/gcx-wrapper.sh datasources clickhouse query "
  SELECT
    file,
    classname,
    name,
    count() AS runs,
    countIf(failure_count = 0 AND error_count = 0 AND skipped_count = 0) AS successful,
    countIf(failure_count > 0 OR error_count > 0) AS fails,
    countIf(skipped_count > 0) AS skipped
  FROM tests.all_test_runs
  WHERE time_inserted >= now() - INTERVAL 7 DAY
    AND file = 'lazy/test_ts_opinfo.py'
  GROUP BY file, classname, name
  ORDER BY runs DESC"

CI Infrastructure

CI infrastructure metrics are stored in grafanacloud-pytorchci-prom. To get a better understanding of the data, clone these repositories in a temporary directory:

Example Queries

Which runner types have the deepest queue right now (jobs assigned but not yet running)?

.claude/skills/ci-metrics/gcx-wrapper.sh datasources prometheus query -d grafanacloud-prom 'topk(10, clamp_min(sum by (name) (gha_assigned_jobs) - sum by (name) (gha_running_jobs), 0))'

How many jobs were running per cluster over the last 6 hours, sampled every 30 minutes? Use --since/--step (or --from/--to) for a range query:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources prometheus query -d grafanacloud-prom 'sum by (cluster) (gha_running_jobs)' --since 6h --step 30m

來自 pytorch 的更多技能

zephyr
pytorch
為嵌入式開發板建置並配置 ExecuTorch 作為 Zephyr RTOS 模組。用於設定包含 ET 的 Zephyr 工作區、新增開發板支援(覆蓋層、…)
aoti-debug
pytorch
調試 AOTInductor (AOTI) 錯誤與崩潰。用於遇到 AOTI 段錯誤、設備不匹配錯誤、常量加載失敗或運行時錯誤時…
skill-writer
pytorch
為 Claude Code 建立結構化 Agent Skills 的指南,包含最佳實踐與驗證方法。涵蓋完整的 Skill 生命週期:範圍界定、檔案結構、YAML 前置資料驗證、內容組織與測試流程。強制執行嚴格的命名規則(小寫、連字號、最多 64 個字元)與描述要求(特定觸發條件、檔案類型、「什麼」與「何時」子句)。提供常見模式的範本,包括唯讀 Skills、基於腳本的 Skills,以及多檔案 Skills 搭配...
triaging-issues
pytorch
根據路由將GitHub問題分派給值班團隊、套用標籤,並關閉提問。適用於處理新的PyTorch問題,或當被要求對某個問題進行分類時…
wheel-size-analyzer
pytorch
使用 GitHub Actions artifacts API 分析 PyTorch 夜間版 wheel 在日期範圍內的大小。用於追蹤二進位檔案大小變化、識別 wheel 大小…
release-go-live-binary-build-matrix
pytorch
當 PyTorch 版本正式發佈時,更新 tools/scripts/generate_binary_build_matrix.py。將 CURRENT_STABLE_VERSION 推進至新的穩定版本,並提升…
pr-review
pytorch
審查 PyTorch 的拉取請求,針對程式碼品質、測試覆蓋率、安全性及向後相容性。適用於審查 PR 時、被要求審查程式碼變更時…
qualcomm
pytorch
建置、測試或開發 QNN(Qualcomm AI Engine Direct)後端。在處理 backends/qualcomm/、建置 QNN(使用…