resilience-program-design

작성자: aws

회복탄력성 프로그램을 설계합니다: 조직, 팀, 또는 포트폴리오 전반에 걸쳐 회복탄력성 정책을 구조화하고 표준화하는 방법(계층적 정책 모델과 함께…)

npx skills add https://github.com/aws/agent-toolkit-for-aws --skill resilience-program-design

Resilience Program Design

Overview

Planning-level guidance for an organization's resilience program: how to structure policies by tier, and how often to run resilience activities.

Structuring resilience policies across an organization

Recommend a tiered policy model (not one policy per service): classify services by business criticality and set policy targets accordingly.

  • Availability SLO — higher as criticality rises. The API accepts only a fixed set of SLO values and rejects out-of-set ones, so confirm the valid values from the API/docs (e.g. aws resiliencehubv2 create-policy help or the Resilience Hub documentation) rather than relying on a hardcoded list — illustratively, values such as 99.9/99.95/99.99.
  • RTO/RPO — tighten as criticality rises (single-digit minutes for critical, hours for low).
  • DR approach — match to criticality (more aggressive for more critical), using a value from the API's DR-approach enum — verify the valid set via the API/docs (e.g. aws resiliencehubv2 create-policy help); illustratively ACTIVE_ACTIVE … BACKUP_AND_RESTORE.

Example (illustrative — resolve the actual enum values against the API before recommending): payments/auth → 99.99 + single-digit-minute RTO + ACTIVE_ACTIVE; internal tools → 99.9 + tens-of-minutes RTO + WARM_STANDBY; dev/test → 99.9 + multi-hour RTO + BACKUP_AND_RESTORE.

Warn against contradictory policies (e.g. the maximum SLO 99.99 with BACKUP_AND_RESTORE, or multi-region RTO shorter than multi-AZ RTO).

How often to run resilience activities (cadence)

Recommend this minimum cadence when asked how often to run resilience activities:

  • Continuous: ARC zonal autoshift practice runs (automated)
  • Weekly: review the Resilience Hub findings dashboard
  • Monthly: run FIS experiments (single-service fault-injection tests)
  • Quarterly: cross-service GameDay
  • Event-driven: after every production incident and before/after major deployments

Security Considerations

Program-level guidance — bake security into the standards you set:

  • Standardize least privilege: require every resilience role (Resilience Hub invoker, FIS execution, ARC operator) in your templates and policies to be least-privilege and resource-scoped, with aws:SourceArn / aws:SourceAccount condition keys on their trust policies to prevent confused-deputy access.
  • Mandate short-lived credentials: require all resilience automation to authenticate as IAM roles with short-lived credentials (role assumption, AWS SSO, instance profiles) — never IAM users with long-lived access keys — as a program standard, since these roles perform privileged and potentially destructive operations.
  • Mandate encryption: make SSE-KMS on report/state buckets part of your tier baseline, and enforce encryption in transit (TLS) — e.g. an aws:SecureTransport deny-if-false condition on those bucket policies and HTTPS-only API access.
  • Govern FIS in production: define an authorization / change-management gate for production fault injection as part of the program cadence.
  • Limit exposure of resilience outputs: assessment findings, FIS logs, and GameDay reports can contain sensitive architectural detail (resource ARNs, IPs, failure modes) — make restricting their access to authorized personnel part of your program standards.
  • Further reading: point teams to the AWS Well-Architected Security Pillar, FIS Security Best Practices, and IAM Best Practices for implementing these standards.

aws의 다른 스킬

analyzing-release-readiness
aws
GitHub PR, GitLab MR 또는 로컬 브랜치에서 병합 전 릴리스 준비 검토를 트리거합니다. 사용자가 코드 변경 사항의 위험성, 정확성 등을 분석하려 할 때 사용합니다.
scanning-with-aws-security-agent
aws
작업 공간에서 AWS Security Agent 스캔 실행 — 소스를 AWS에 업로드하고, 관리형 Security Agent 서비스로 스캔한 후, 순위가 매겨진 검증된 결과를 반환합니다…
coordinating-multi-space-devops-agent
aws
하나의 Claude Code 세션에서 여러 AgentSpaces에 걸쳐 AWS DevOps Agent를 조정하세요 — 질문을 올바른 공간(프로덕션 vs 스테이징 vs 지식)으로 라우팅하고,…
aws-security
aws
AWS 보안 서비스 및 워크플로우를 다룹니다 — Security Hub V2 (OCSF) findings, 커넥터, 애그리게이터, 자동화 규칙, 보안 상태 요약 등…
querying-aws-sagemaker-catalog
aws
SageMaker Catalog 자산 메타데이터 테이블에서 SQL 분석을 실행하며, S3 Tables에서 Apache Iceberg로 내보낸 데이터를 대상으로 합니다. 거버넌스 쿼리, 자산 성장 추적 등을 다룹니다.
agents-connect
aws
에이전트를 Gateway를 통해 외부 API, 도구 또는 서비스에 연결하거나 Cedar 정책으로 도구 접근을 제한할 때 사용합니다. 게이트웨이 설정, 대상...
aurora-dsql
aws
Aurora DSQL 클러스터를 프로비저닝하고 관리하며, psql 또는 DSQL 커넥터를 통해 연결하고, 스키마를 관리하고, 쿼리를 실행하고, MySQL에서 마이그레이션하고, 쿼리 계획을 진단합니다.
transitgateway
aws
AWS Transit Gateway를 구성합니다: 허브를 생성하고 VPC를 연결하며, 라우팅 테이블로 트래픽을 분리하고, 허브를 통해 이그레스 및 검사를 중앙화합니다…