ci-metrics

PyTorch CI, GitHub Actions, HUD, Grafana und Infrastrukturmetriken abfragen. Verwenden, wenn Benutzer nach CI-Dauer, Jobfehlern, Wartezeiten, Workflow-Trends fragen…

npx skills add https://github.com/pytorch/pytorch --skill ci-metrics

PyTorch CI Metrics

PyTorch CI and infrastructure metrics are exposed through Grafana. Use .claude/skills/ci-metrics/gcx-wrapper.sh for all Grafana access; it configures the PyTorch Grafana server, context, and authentication. Only users with write permission to the repo have access to Grafana. The authentication only provides read only access.

Requirements

The wrapper needs these tools on PATH:

  • gh - fetches the Grafana token and must be authenticated; if not, run gh auth login --hostname github.com --git-protocol ssh --web.
  • curl - downloads gcx and fetches the token from HUD.

On first use the wrapper downloads a pinned, checksum-verified gcx binary into a private cache (~/.cache/pytorch-ci-metrics/) and authenticates automatically. Nothing is installed on your PATH. If a tool is missing or gh is not authenticated, it exits with an error describing what to fix.

Datasources

Get the list of datasources available:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources list

The data contains metrics for many repos owned by the PyTorch repo. When possible, restrict queries to just the pytorch/pytorch repository.

CI and Test Run Data

CI and test run data are stored in grafana-clickhouse-datasource. List all the available tables:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources clickhouse list-tables

Important dataset:

  • GitHub webhook data
  • Tests
    • Database: tests
    • tests.all_test_runs - contains every test run. This is an extremely large table, so be considerate with filtering and timing.
    • Do not use tests.test_run_s3 as it contains partial data only.

To get additional guidance on common queries, clone https://github.com/pytorch/test-infra into a temporary directory and read the torchci folder.

Example Queries

Within the pytorch/pytorch repo on main, list the top most failing workflow jobs in the last 2 weeks:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources clickhouse query "
  SELECT name, count(DISTINCT id) AS failures
  FROM default.workflow_job
  WHERE conclusion = 'failure'
    AND completed_at >= now() - INTERVAL 2 WEEK
    AND repository_full_name = 'pytorch/pytorch'
    AND head_branch = 'main'
  GROUP BY name ORDER BY failures DESC LIMIT 10"

For a test file, how many times was it run in the last week? How many times did it pass or fail?

.claude/skills/ci-metrics/gcx-wrapper.sh datasources clickhouse query "
  SELECT
    file,
    classname,
    name,
    count() AS runs,
    countIf(failure_count = 0 AND error_count = 0 AND skipped_count = 0) AS successful,
    countIf(failure_count > 0 OR error_count > 0) AS fails,
    countIf(skipped_count > 0) AS skipped
  FROM tests.all_test_runs
  WHERE time_inserted >= now() - INTERVAL 7 DAY
    AND file = 'lazy/test_ts_opinfo.py'
  GROUP BY file, classname, name
  ORDER BY runs DESC"

CI Infrastructure

CI infrastructure metrics are stored in grafanacloud-pytorchci-prom. To get a better understanding of the data, clone these repositories in a temporary directory:

Example Queries

Which runner types have the deepest queue right now (jobs assigned but not yet running)?

.claude/skills/ci-metrics/gcx-wrapper.sh datasources prometheus query -d grafanacloud-prom 'topk(10, clamp_min(sum by (name) (gha_assigned_jobs) - sum by (name) (gha_running_jobs), 0))'

How many jobs were running per cluster over the last 6 hours, sampled every 30 minutes? Use --since/--step (or --from/--to) for a range query:

.claude/skills/ci-metrics/gcx-wrapper.sh datasources prometheus query -d grafanacloud-prom 'sum by (cluster) (gha_running_jobs)' --since 6h --step 30m

Mehr Skills von pytorch

zephyr
pytorch
Erstelle und konfiguriere ExecuTorch als Zephyr RTOS-Modul für eingebettete Boards. Verwende beim Einrichten eines Zephyr-Workspace mit ET, Hinzufügen von Board-Unterstützung (Overlays,…
aoti-debug
pytorch
Debuggen von AOTInductor (AOTI)-Fehlern und Abstürzen. Verwenden bei AOTI-Segmentierungsfehlern, Gerätekonflikten, Fehlern beim Laden von Konstanten oder Laufzeitfehlern von…
skill-writer
pytorch
We need to translate the given English text into German, preserving the name "skill-writer" if it appears. The instruction says: "Do not include the name unless it appears in the source text." The source text does not contain "skill-writer" explicitly. The name to preserve is "skill-writer" but it's not in the text. So we just translate the text. The text: "Guide for creating well-structured Agent Skills for Claude Code with best practices and validation. Covers full Skill lifecycle: scoping, file structure, YAML frontmatter validation, content organization, and testing procedures Enforces strict naming rules (lowercase, hyphens, max 64 chars) and description requirements (specific triggers, file types, "what" and "when" clauses) Provides templates for common patterns including read-only Skills, script-based Skills, and multi-file Skills with..." We need to translate accurately, preserving product names like "Claude Code", "YAML", technical terms, numbers, etc. Also note the ellipsis at the end. Translation: "Le
triaging-issues
pytorch
Leitet GitHub-Issues weiter, indem es sie an Bereitschaftsteams weiterleitet, Labels anwendet und Fragen schließt. Verwenden Sie dies bei der Verarbeitung neuer PyTorch-Issues oder wenn Sie aufgefordert werden, ein Issue zu triagieren…
wheel-size-analyzer
pytorch
Analysiere die Größen nächtlicher PyTorch-Wheel-Builds über einen Datumsbereich mithilfe der GitHub-Actions-Artefakte-API. Verwende dies zur Verfolgung von Änderungen der Binärgröße, zur Identifizierung von Wheel-Größen…
release-go-live-binary-build-matrix
pytorch
Aktualisiert tools/scripts/generate_binary_build_matrix.py, wenn ein PyTorch-Release live geht. Erhöht CURRENT_STABLE_VERSION auf die neue stabile Version, befördert die…
pr-review
pytorch
Überprüfe PyTorch-Pull-Requests auf Codequalität, Testabdeckung, Sicherheit und Rückwärtskompatibilität. Verwende dies beim Überprüfen von PRs, wenn du gebeten wirst, Codeänderungen zu überprüfen,…
qualcomm
pytorch
Erstellen, testen oder entwickeln Sie das QNN (Qualcomm AI Engine Direct) Backend. Verwenden Sie, wenn Sie an backends/qualcomm/ arbeiten, QNN erstellen (verwenden…