migrate-workflow-ec2-to-osdc

Schritt-für-Schritt-Playbook zur Migration eines pytorch/pytorch .github/workflows/*.yml von EC2 zu OSDC (ARC)-Runnern – deckt sowohl Dial-up- als auch 100% Opt-in-Muster ab,…

npx skills add https://github.com/pytorch/test-infra --skill migrate-workflow-ec2-to-osdc

Migrating a workflow from EC2 to OSDC

OSDC (= ARC = on-site data center, EKS-hosted self-hosted runners) replaces EC2 runners. Migration touches the workflow file and requires a few inputs to flow into the reusable _linux-build.yml / _linux-test.yml so the OSDC code paths (build-osdc, test-osdc) activate. EC2 paths (build, test) and OSDC paths gate on inputs.use-arc (!inputs.use-arc vs inputs.use-arc), so flipping that input swaps execution lanes.

This playbook is for workflows that call the reusable _linux-build.yml / _linux-test.yml (e.g. test-b200.yml, pull.yml, trunk.yml, tsan.yml, operator_microbenchmark.yml). Standalone OSDC workflows (raw ARC label + container: directive) follow a different pattern not covered here.

Decide: dial-up vs. 100% opt-in

PatternWhen to useExample
Dial-up (preferred)Existing workflow with broad coverage; you want to ramp OSDC adoption via labels like pull.yml/trunk.yml.test-b200.yml (PR #181544)
100% opt-inWorkflow is meant to always run on OSDC (e.g. testing the single B200 we own on EKS).operator_microbenchmark.yml, attention_op_microbenchmark.yml (B200 jobs)

The two pieces — runner_prefix and use-arc — must move together. Hardcoding one but driving the other off the determinator is a bug.

Migration steps (dial-up pattern)

Worked reference: PR #181544 / commit f156b7ddfd1 ("Migrate smoke test on B200 to OSDC"). The diff was 10 lines.

1. Make sure get-label-type opts into the ARC experiment

In the get-label-type job that calls _runner-determinator.yml, add:

check_experiments: arc,lf

Without this the determinator won't consider the ARC experiment and use-arc will always be false.

2. Plumb five inputs into the build job

On the call to ./.github/workflows/_linux-build.yml:

with:
  runner_prefix: "${{ needs.get-label-type.outputs.label-type }}"  # likely already present
  ...existing inputs...
  ci-docker-hash: ${{ needs.get-label-type.outputs.ci-docker-hash }}
  use-arc: ${{ needs.get-label-type.outputs.use-arc == 'true' }}
  python-version: "3.10"     # match the docker-image-name's python
  compiler: gcc11            # match the docker-image-name's compiler
  cuda-version: "13.0"       # match the docker-image-name's cuda (or "cpu" for CPU)

ci-docker-hash is the git rev-parse HEAD:.ci/docker hash that the determinator computes. _linux-build.yml appends it to the ghcr.io image tag (ghcr.io/pytorch/<docker-image-name>-<hash>); without it the OSDC build pulls the unhashed tag, which may not exist or may resolve to a stale image. The build job must have needs: get-label-type for this expression to resolve.

The last three (python-version / compiler / cuda-version) feed setup-linux so it can configure the OSDC container env. Read them off the existing docker-image-name (e.g. pytorch-linux-jammy-cuda13.0-cudnn9-py3-gcc11 → py3.10 / gcc11 / cuda13.0).

3. Add get-label-type to the test job's needs and plumb the use-arc + setup-linux inputs

On the call to ./.github/workflows/_linux-test.yml:

needs:
  - <existing-build-job>
  - get-label-type           # add this
with:
  ...existing inputs...
  use-arc: ${{ needs.get-label-type.outputs.use-arc == 'true' }}
  python-version: "3.10"
  compiler: gcc11
  cuda-version: "13.0"

The test job needs a direct needs: get-label-type dependency to read its outputs (it can't rely on transitive needs through the build job). ci-docker-hash is not an input to _linux-test.yml — the test job picks the right image up via needs.<build>.outputs.docker-image.

4. Drop the OSDC-incompatible inputs

If the EC2 path used inputs that don't apply on OSDC (e.g. aws-role-to-assume: for ECR pulls), remove them. OSDC has its own AWS role wired into setup-linux (arn:aws:iam::308535385114:role/arc) and pulls images from ghcr.io/pytorch, not ECR.

In #181544 this meant deleting:

aws-role-to-assume: arn:aws:iam::308535385114:role/gha_workflow_s3_and_ecr_read_only

5. Verify

  • gh workflow run or push a PR with the relevant ciflow/* tag.
  • Look for the build-osdc / test-osdc jobs running (they're separate jobs in _linux-build.yml / _linux-test.yml, gated on inputs.use-arc). When use-arc is false, the original build / test jobs run instead.

Migration steps (100% opt-in pattern)

For a job that should always run on OSDC, hardcode runner_prefix and use-arc instead of pulling them from the determinator. The build job still needs ci-docker-hash from get-label-type, so keep needs: get-label-type:

needs: get-label-type
with:
  runner_prefix: "mt-"
  ci-docker-hash: ${{ needs.get-label-type.outputs.ci-docker-hash }}
  use-arc: true
  python-version: "3.10"
  compiler: gcc11
  cuda-version: "13.0"

Both jobs (build and test) need the use-arc + python-version / compiler / cuda-version inputs; only the build job needs ci-docker-hash. The build job's runner: stays as the EC2 label (e.g. linux.r7i.4xlarge); _linux-build.yml's build-osdc job translates that to the right ARC runner via map_ec2_to_arc.py.

Reference: operator_microbenchmark.yml jobs opmicrobenchmark-build-b200 / opmicrobenchmark-test-b200.

What the inputs do (mental model)

_linux-build.yml and _linux-test.yml each define two jobs:

  • build / test — EC2 path. Runs directly on the runner host VM. Gated on !inputs.use-arc.
  • build-osdc / test-osdc — OSDC path. Runs inside a container: from ghcr.io/pytorch/${docker-image-name}${ci-docker-hash ? '-' + ci-docker-hash : ''}. Gated on inputs.use-arc.

The OSDC jobs call setup-linux with use-arc: true and pass python-version / compiler / cuda-version so the container env matches the build environment. That's why the migration must pass these three inputs — setup-linux errors out without them on the OSDC path. ci-docker-hash is what pins the ghcr.io image tag to a specific .ci/docker tree state so the build doesn't drift onto an unhashed (or wrong) image.

Common gotchas

  • Forgetting check_experiments: arc,lf → use-arc is always false, so OSDC never activates and you'll think the migration silently failed.
  • Forgetting needs: get-label-type on the test job → ${{ needs.get-label-type.outputs.use-arc }} evaluates to empty, dial-up doesn't engage.
  • Forgetting ci-docker-hash on the build job → OSDC pulls the unhashed ghcr.io tag (<docker-image-name> with no -<hash> suffix), which may not exist or may resolve to a stale image. The failure mode is a manifest unknown pull error in the OSDC container step.
  • Plumbing ci-docker-hash from needs.get-label-type.outputs.* without adding needs: get-label-type → the expression resolves to empty, so it's a silent no-op equivalent to not setting it at all.
  • Mismatched python-version / compiler / cuda-version vs. docker-image-name → container has one toolchain, setup-linux configures another, build fails confusingly.
  • Leaving aws-role-to-assume: for ECR → harmless on OSDC (it's only used by EC2 path) but stale and misleading; remove it.
  • Mixing patterns (dynamic runner_prefix + hardcoded use-arc: true, or vice versa). Pick one pattern and apply it consistently.

Mehr Skills von pytorch

zephyr
pytorch
Erstelle und konfiguriere ExecuTorch als Zephyr RTOS-Modul für eingebettete Boards. Verwende beim Einrichten eines Zephyr-Workspace mit ET, Hinzufügen von Board-Unterstützung (Overlays,…
aoti-debug
pytorch
Debuggen von AOTInductor (AOTI)-Fehlern und Abstürzen. Verwenden bei AOTI-Segmentierungsfehlern, Gerätekonflikten, Fehlern beim Laden von Konstanten oder Laufzeitfehlern von…
skill-writer
pytorch
We need to translate the given English text into German, preserving the name "skill-writer" if it appears. The instruction says: "Do not include the name unless it appears in the source text." The source text does not contain "skill-writer" explicitly. The name to preserve is "skill-writer" but it's not in the text. So we just translate the text. The text: "Guide for creating well-structured Agent Skills for Claude Code with best practices and validation. Covers full Skill lifecycle: scoping, file structure, YAML frontmatter validation, content organization, and testing procedures Enforces strict naming rules (lowercase, hyphens, max 64 chars) and description requirements (specific triggers, file types, "what" and "when" clauses) Provides templates for common patterns including read-only Skills, script-based Skills, and multi-file Skills with..." We need to translate accurately, preserving product names like "Claude Code", "YAML", technical terms, numbers, etc. Also note the ellipsis at the end. Translation: "Le
triaging-issues
pytorch
Leitet GitHub-Issues weiter, indem es sie an Bereitschaftsteams weiterleitet, Labels anwendet und Fragen schließt. Verwenden Sie dies bei der Verarbeitung neuer PyTorch-Issues oder wenn Sie aufgefordert werden, ein Issue zu triagieren…
wheel-size-analyzer
pytorch
Analysiere die Größen nächtlicher PyTorch-Wheel-Builds über einen Datumsbereich mithilfe der GitHub-Actions-Artefakte-API. Verwende dies zur Verfolgung von Änderungen der Binärgröße, zur Identifizierung von Wheel-Größen…
release-go-live-binary-build-matrix
pytorch
Aktualisiert tools/scripts/generate_binary_build_matrix.py, wenn ein PyTorch-Release live geht. Erhöht CURRENT_STABLE_VERSION auf die neue stabile Version, befördert die…
pr-review
pytorch
Überprüfe PyTorch-Pull-Requests auf Codequalität, Testabdeckung, Sicherheit und Rückwärtskompatibilität. Verwende dies beim Überprüfen von PRs, wenn du gebeten wirst, Codeänderungen zu überprüfen,…
qualcomm
pytorch
Erstellen, testen oder entwickeln Sie das QNN (Qualcomm AI Engine Direct) Backend. Verwenden Sie, wenn Sie an backends/qualcomm/ arbeiten, QNN erstellen (verwenden…