ci-metrics

作者: pytorch

查询PyTorch CI、GitHub Actions、HUD、Grafana和基础设施指标。当用户询问CI持续时间、任务失败、队列时间、工作流趋势等时使用。

npx skills add https://github.com/pytorch/pytorch --skill ci-metrics

PyTorch CI Metrics

PyTorch CI and infrastructure metrics are exposed through Grafana. Use .claude/skills/ci-metrics/gcx-wrapper.sh for all Grafana access; it configures the PyTorch Grafana server, context, and authentication. Only users with write permission to the repo have access to Grafana. The authentication only provides read only access.

Requirements

The wrapper needs these tools on PATH:

  • gh - fetches the Grafana token and must be authenticated; if not, run gh auth login --hostname github.com --git-protocol ssh --web.
  • curl - downloads gcx and fetches the token from HUD.

On first use the wrapper downloads a pinned, checksum-verified gcx binary into a private cache (~/.cache/pytorch-ci-metrics/) and authenticates automatically. Nothing is installed on your PATH. If a tool is missing or gh is not authenticated, it exits with an error describing what to fix.

Datasources

Get the list of datasources available:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources list

The data contains metrics for many repos owned by the PyTorch repo. When possible, restrict queries to just the pytorch/pytorch repository.

CI and Test Run Data

CI and test run data are stored in grafana-clickhouse-datasource. List all the available tables:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources clickhouse list-tables

Important dataset:

  • GitHub webhook data
  • Tests
    • Database: tests
    • tests.all_test_runs - contains every test run. This is an extremely large table, so be considerate with filtering and timing.
    • Do not use tests.test_run_s3 as it contains partial data only.

To get additional guidance on common queries, clone https://github.com/pytorch/test-infra into a temporary directory and read the torchci folder.

Example Queries

Within the pytorch/pytorch repo on main, list the top most failing workflow jobs in the last 2 weeks:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources clickhouse query "
  SELECT name, count(DISTINCT id) AS failures
  FROM default.workflow_job
  WHERE conclusion = 'failure'
    AND completed_at >= now() - INTERVAL 2 WEEK
    AND repository_full_name = 'pytorch/pytorch'
    AND head_branch = 'main'
  GROUP BY name ORDER BY failures DESC LIMIT 10"

For a test file, how many times was it run in the last week? How many times did it pass or fail?

.claude/skills/ci-metrics/gcx-wrapper.sh datasources clickhouse query "
  SELECT
    file,
    classname,
    name,
    count() AS runs,
    countIf(failure_count = 0 AND error_count = 0 AND skipped_count = 0) AS successful,
    countIf(failure_count > 0 OR error_count > 0) AS fails,
    countIf(skipped_count > 0) AS skipped
  FROM tests.all_test_runs
  WHERE time_inserted >= now() - INTERVAL 7 DAY
    AND file = 'lazy/test_ts_opinfo.py'
  GROUP BY file, classname, name
  ORDER BY runs DESC"

CI Infrastructure

CI infrastructure metrics are stored in grafanacloud-pytorchci-prom. To get a better understanding of the data, clone these repositories in a temporary directory:

Example Queries

Which runner types have the deepest queue right now (jobs assigned but not yet running)?

.claude/skills/ci-metrics/gcx-wrapper.sh datasources prometheus query -d grafanacloud-prom 'topk(10, clamp_min(sum by (name) (gha_assigned_jobs) - sum by (name) (gha_running_jobs), 0))'

How many jobs were running per cluster over the last 6 hours, sampled every 30 minutes? Use --since/--step (or --from/--to) for a range query:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources prometheus query -d grafanacloud-prom 'sum by (cluster) (gha_running_jobs)' --since 6h --step 30m

来自 pytorch 的更多技能

zephyr
pytorch
将ExecuTorch构建并配置为嵌入式板卡的Zephyr RTOS模块。用于设置包含ET的Zephyr工作区、添加板级支持(覆盖层、…)时。
aoti-debug
pytorch
调试AOTInductor(AOTI)的错误和崩溃。在遇到AOTI段错误、设备不匹配错误、常量加载失败或运行时错误时使用…
skill-writer
pytorch
为Claude Code创建结构良好的Agent Skills指南,包含最佳实践与验证方法。涵盖Skill完整生命周期:范围界定、文件结构、YAML前置元数据验证、内容组织及测试流程。强制执行严格命名规则(小写字母、连字符、最长64字符)及描述要求(具体触发条件、文件类型、“什么”和“何时”子句)。提供常见模式模板,包括只读型Skills、脚本型Skills及多文件型Skills等。
triaging-issues
pytorch
通过路由到值班团队、应用标签和关闭问题来分类GitHub问题。在处理新的PyTorch问题或被要求分类时使用…
wheel-size-analyzer
pytorch
使用 GitHub Actions artifacts API 分析指定日期范围内 PyTorch 夜间版 wheel 的大小。用于跟踪二进制大小变化、识别 wheel 大小…
release-go-live-binary-build-matrix
pytorch
当PyTorch版本正式发布时,更新tools/scripts/generate_binary_build_matrix.py。将CURRENT_STABLE_VERSION推进到新的稳定版本,并提升…
pr-review
pytorch
审查PyTorch拉取请求的代码质量、测试覆盖率、安全性和向后兼容性。在审查PR、被要求审查代码变更时使用…
qualcomm
pytorch
构建、测试或开发QNN(高通AI引擎直连)后端。在处理backends/qualcomm/目录、构建QNN(使用……时使用。