hf-cloud-sagemaker-iam-preflight

作者: huggingface

在部署或訓練之前,確保存在一個可用的 SageMaker 執行角色。每當要建立 SageMaker 端點、模型、訓練…時,請使用此技能。

npx skills add https://github.com/huggingface/skills --skill hf-cloud-sagemaker-iam-preflight

SageMaker IAM Preflight

Every SageMaker resource needs an execution role — the IAM role SageMaker assumes to read model artifacts from S3, pull serving containers from ECR, and write logs. Most deployments fail here because the script tried to create a new role without checking if a usable one already existed, then blew up because the caller is an SSO principal.

This skill encodes the right order: discover, validate, only create if necessary.

Running the helpers (cross-platform)

The helpers are Python so they run identically on Windows, macOS, and Linux:

python3 scripts/check_role.py        # macOS / Linux
python  scripts/check_role.py        # Windows (PowerShell / cmd)

Run them from the shell where the AWS CLI already works — i.e. wherever aws sts get-caller-identity succeeds. The script shells out to that same aws binary and inherits the shell's profile, region, SSO session, proxy, and credential chain.

Windows / WSL / Git Bash caveat. Do not invoke these through a Bash shim (WSL, Git Bash, MSYS) on Windows. Those Bash environments frequently do not share the Windows AWS config, credentials, SSO sessions, environment variables, or proxy settings — so aws sts get-caller-identity fails inside Bash even when it works natively in PowerShell. (This is exactly why the old .sh helpers failed on Windows and were replaced with Python.) If you're in PowerShell, run python ...\check_role.py directly in PowerShell. If the helper still can't see your identity, run the same discovery natively (see "Native AWS CLI equivalent" below) in the shell where aws sts get-caller-identity returns your ARN.

Order of operations

Step 1 — Did the user provide a role?

Validate that one specifically:

python3 scripts/check_role.py "<role-name-or-arn>"

On success it prints the ARN to stdout (exit 0). On failure it logs why on stderr. Don't try to silently fix a broken role — surface the problem.

Step 2 — Discover existing roles

python3 scripts/check_role.py

Lists roles matching common SageMaker patterns (AmazonSageMaker-ExecutionRole-*, SageMakerExecutionRole*, etc.), ranks by last-used date (most recent first), validates trust policy in that order, returns the first usable ARN. Most accounts that have used SageMaker before already have one.

Why rank by last-used: in accounts with multiple roles (auto-generated 2021 role + manual project role + etc.), the alphabetically-first one is rarely the actively-maintained one. The most-recently-used role is more likely to have current policies — including cross-account ECR pull. The script prints the ranking so you can see which got picked.

IAM frequently reports no RoleLastUsed at all (tracking only covers recent activity). When every candidate ties at "never used", the script falls back to newest creation date — a newer role is more likely to have current policies than a 2021 leftover.

Step 3 — Create, only if discovery found nothing

If the user can create (has IAM permissions):

python3 scripts/create_role.py "<role-name>" "<model-bucket>"

Second arg scopes S3 access to a specific bucket. Omit if unknown; script warns and the user can update the policy later.

If the user cannot create (SSO principal — hf-cloud-aws-context-discovery will have flagged this):

Stop and surface this clearly. Don't retry alternative IAM operations hoping one works:

I can't find an existing SageMaker execution role, and you're authenticated via SSO so you can't create one directly. Please either:

  • Ask your AWS admin for a SageMaker execution role ARN, or
  • Have them grant your SSO permission set iam:CreateRole, iam:PutRolePolicy

Specific instructions get unblocked fast; vague "permission denied" messages don't.

What "validated" means

A role is usable when (1) it exists, (2) its trust policy allows sagemaker.amazonaws.com to sts:AssumeRole, and (3) its permissions grant only the actions and resources this deployment needs. See references/trust-policy.json for the canonical trust policy.

check_role.py verifies existence and trust because policy evaluation depends on the deployment's exact S3, ECR, logging, and optional output resources. Before deployment, inspect the selected role's policies and compare them with references/minimum-permissions.json; add only missing actions and scope them to the required resources. Do not attach AmazonSageMakerFullAccess or defer permission review until an AccessDenied failure.

Minimum permissions

references/minimum-permissions.json is the standalone inline policy for endpoint execution:

  • s3:GetObject + s3:ListBucket on the model artifact bucket
  • ECR pull permissions
  • CloudWatch logs and metrics

create_role.py installs this inline policy without attaching a managed FullAccess policy. Replace REPLACE_WITH_MODEL_BUCKET in the template with the actual bucket name — create_role.py does this automatically when given a bucket as its second argument. Add narrowly scoped permissions separately for optional features such as async output or data capture.

Native AWS CLI equivalent (fallback)

If the Python helper can't run or can't see your identity (rare — usually a broken PATH or running under a Bash shim that lacks AWS context), do the same preflight by hand in the shell where aws sts get-caller-identity works. The logic is just AWS CLI calls; the helper exists only to bundle and rank them.

PowerShell:

# 1. List candidate SageMaker roles
aws iam list-roles --query "Roles[?contains(RoleName,'SageMaker') || contains(RoleName,'sagemaker')]" --output json

# 2. For each candidate, confirm the trust policy allows sagemaker.amazonaws.com
aws iam get-role --role-name <role-name> --query "Role.AssumeRolePolicyDocument" --output json

# 3. Prefer the most-recently-used role with SageMaker-execution naming
#    (LastUsedDate is often None for every role — then prefer newest CreateDate)
aws iam get-role --role-name <role-name> --query "Role.[RoleLastUsed.LastUsedDate, CreateDate]" --output text

Pick the most-recently-used role whose trust policy contains sagemaker.amazonaws.com. Use the resulting ARN exactly as if check_role.py had returned it. Bash/macOS/Linux use the same commands.

來自 huggingface 的更多技能

sync-models
huggingface
將 chat-ui 的模型配置與 HuggingFace 路由器同步——為新模型新增描述、標記具備推理能力的模型、為 32B 以上的模型啟用 artifacts…
custom-blocks
huggingface
Use when the user has written (or wants to write) a `ModularPipelineBlocks` subclass in a local Python file and needs to package it into a Hub-uploadable…
self-review
huggingface
Use before opening a PR, or whenever asked to self-review a diffusers contribution. Applies the same rubric as the `@claude` CI (checks the diff against…
hf-cloud-sagemaker-production-defaults
huggingface
建立一個啟用自動擴展、CloudWatch 警報和標籤功能的 SageMaker 端點(即時或非同步)。每當即將建立…時,請使用此技能。
hf-cloud-serving-image-selection
huggingface
為 SageMaker 模型部署選擇合適的服務容器,並找出其目前的映像 URI。每當要將模型部署到…時,請使用此技能。
Hugging Face Cli
huggingface
Execute Hugging Face Hub operations using the `hf` CLI. Use when the user needs to download models/datasets/spaces, upload files to Hub repositories, create repos, manage local cache, or run compute jobs on HF infrastructure. Covers authentication, file transfers, repository creation, cache operations, and cloud compute.
Hugging Face Datasets
huggingface
在 Hugging Face Hub 上建立與管理資料集。支援初始化儲存庫、定義配置/系統提示、串流更新資料列,以及基於 SQL 的資料集查詢/轉換。設計與 HF MCP 伺服器搭配使用,以實現完整的資料集工作流程。
Hugging Face Evaluation
huggingface
在 Hugging Face 模型卡中新增與管理評估結果。支援從 README 內容中提取評估表格、從 Artificial Analysis API 匯入分數,以及使用 vLLM/lighteval 執行自訂模型評估。可搭配 model-index 中繼資料格式使用。