aws-lambda-durable-functions

작성자: aws

AWS Lambda 지속 함수를 사용하여 자동 상태 유지, 재시도 로직, 오케스트레이션을 갖춘 탄력적이고 장기 실행되는 다단계 애플리케이션을 구축합니다…

npx skills add https://github.com/aws/agent-toolkit-for-aws --skill aws-lambda-durable-functions

AWS Lambda durable functions

Build resilient multi-step applications and AI workflows that can execute for up to 1 year while maintaining reliable progress despite interruptions.

Works best with the AWS MCP server but is not required. All AWS interactions in this skill use standard AWS CLI commands that work in any environment with configured AWS credentials.

Critical Rules

Read these before writing any code. Each one is a constraint that will silently break a function if violated.

  1. Durable execution must be enabled at function creation time — it cannot be retrofitted. A new Lambda function must be created with durable execution turned on. Migrate the logic into the new function; do not attempt to install the SDK and wrap the handler of the existing function and expect it to work.
  2. Durable functions must be invoked with a qualified ARN — a specific version, an alias, or the literal $LATEST suffix. An unqualified function name will fail. See the Invocation Requirements section below for examples.
  3. Durable operations cannot be nested. You cannot call context.step(), context.wait(), or context.invoke() from inside another step's callback. Use context.runInChildContext() to group operations instead.
  4. All non-deterministic code must run inside steps. Date.now(), Math.random(), UUID generation, API calls, and database queries outside a step will produce different values on replay and corrupt execution state.
  5. Closure mutations are lost on replay - return values from steps
  6. Side effects outside steps repeat - use context.logger (replay-aware)

When to Load Reference Files

Load the appropriate reference file based on what the user is working on:

  • Getting started, basic setup, example, ESLint, or Jest setup -> see getting-started.md
  • Understanding replay model, determinism, or non-deterministic errors -> see replay-model-rules.md
  • Creating steps, atomic operations, or retry logic -> see step-operations.md
  • Waiting, delays, callbacks, external systems, or polling -> see wait-operations.md
  • Parallel execution, map operations, batch processing, or concurrency -> see concurrent-operations.md
  • Error handling, retry strategies, saga pattern, or compensating transactions -> see error-handling.md
  • Advanced error handling, timeout handling, circuit breakers, or conditional retries -> see advanced-error-handling.md
  • Testing, local testing, cloud testing, test runner, or flaky tests -> see testing-patterns.md
  • Deployment, CloudFormation, CDK, SAM, log groups, deploy, or infrastructure -> see deployment-iac.md
  • Advanced patterns, GenAI agents, completion policies, step semantics, or custom serialization -> see advanced-patterns.md
  • troubleshooting, stuck execution, failed execution, debug execution ID, execution history, execution error, why did my execution fail, execution timed out, callback not received, diagnose execution, or root cause execution -> see troubleshooting-executions.md

Quick Reference

Basic Handler Pattern

TypeScript:

import { withDurableExecution, DurableContext } from '@aws/durable-execution-sdk-js';

export const handler = withDurableExecution(async (event, context: DurableContext) => {
  const result = await context.step('process', async () => processData(event));
  return result;
});

Python:

from aws_durable_execution_sdk_python import durable_execution, DurableContext

@durable_execution
def handler(event: dict, context: DurableContext) -> dict:
    result = context.step(lambda _: process_data(event), name='process')
    return result

Python API Differences

The Python SDK differs from TypeScript in several key areas:

  • Steps: Use @durable_step decorator + context.step(my_step(args)), or inline context.step(lambda _: ..., name='...'). Prefer the decorator for automatic step naming.
  • Wait: context.wait(duration=Duration.from_seconds(n), name='...')
  • Exceptions: ExecutionError (permanent), InvocationError (transient), CallbackError (callback failures)
  • Testing: Use DurableFunctionTestRunner class directly - instantiate with handler, use context manager, call run(input=...)

Invocation Requirements

Durable functions require qualified ARNs (version, alias, or $LATEST):

# Valid
aws lambda invoke --function-name my-function:1 output.json
aws lambda invoke --function-name my-function:live output.json

# Invalid - will fail
aws lambda invoke --function-name my-function output.json

IAM Permissions

Your Lambda execution role MUST have the AWSLambdaBasicDurableExecutionRolePolicy managed policy attached. This includes:

  • lambda:CheckpointDurableExecution - Persist execution state
  • lambda:GetDurableExecutionState - Retrieve execution state
  • CloudWatch Logs permissions

Additional permissions needed for:

  • Durable invokes: lambda:InvokeFunction on target function ARNs
  • External callbacks: Systems need lambda:SendDurableExecutionCallbackSuccess and lambda:SendDurableExecutionCallbackFailure

Validation Guidelines

When writing or reviewing durable function code, ALWAYS check for these replay model violations:

  1. Non-deterministic code outside steps: Date.now(), Math.random(), UUID generation, API calls, database queries must all be inside steps
  2. Nested durable operations in step functions: Cannot call context.step(), context.wait(), or context.invoke() inside a step function — use context.runInChildContext() instead
  3. Closure mutations that won't persist: Variables mutated inside steps are NOT preserved across replays — return values from steps instead
  4. Side effects outside steps that repeat on replay: Use context.logger for logging (it is replay-aware and deduplicates automatically)

When implementing or modifying tests for durable functions, ALWAYS verify:

  1. All operations have descriptive names
  2. Tests get operations by NAME, never by index
  3. Replay behavior is tested with multiple invocations
  4. Use LocalDurableTestRunner for local testing

Security Considerations

  • Checkpoint data encryption: Execution state is persisted automatically. Enable KMS encryption on associated CloudWatch Log Groups to protect checkpointed data at rest.
  • Sensitive data in step results: Step return values are checkpointed and persisted. Do not return secrets, raw credentials, or PII from steps — store sensitive data in Secrets Manager or SSM Parameter Store and return references instead.
  • Input validation: Validate and sanitize event payloads at the handler entry point before passing data to steps.
  • Credential management: Retrieve secrets from AWS Secrets Manager or SSM Parameter Store within steps.
  • Callback payload validation: Data received via waitForCallback originates from external systems — validate and sanitize before processing.
  • Logging: Avoid DEBUG log level in non-development environments as it may expose step results and execution state. Enable CloudWatch Logs encryption with KMS.

Resources

aws의 다른 스킬

analyzing-release-readiness
aws
GitHub PR, GitLab MR 또는 로컬 브랜치에서 병합 전 릴리스 준비 검토를 트리거합니다. 사용자가 코드 변경 사항의 위험성, 정확성 등을 분석하려 할 때 사용합니다.
scanning-with-aws-security-agent
aws
작업 공간에서 AWS Security Agent 스캔 실행 — 소스를 AWS에 업로드하고, 관리형 Security Agent 서비스로 스캔한 후, 순위가 매겨진 검증된 결과를 반환합니다…
coordinating-multi-space-devops-agent
aws
하나의 Claude Code 세션에서 여러 AgentSpaces에 걸쳐 AWS DevOps Agent를 조정하세요 — 질문을 올바른 공간(프로덕션 vs 스테이징 vs 지식)으로 라우팅하고,…
aws-security
aws
AWS 보안 서비스 및 워크플로우를 다룹니다 — Security Hub V2 (OCSF) findings, 커넥터, 애그리게이터, 자동화 규칙, 보안 상태 요약 등…
querying-aws-sagemaker-catalog
aws
SageMaker Catalog 자산 메타데이터 테이블에서 SQL 분석을 실행하며, S3 Tables에서 Apache Iceberg로 내보낸 데이터를 대상으로 합니다. 거버넌스 쿼리, 자산 성장 추적 등을 다룹니다.
agents-connect
aws
에이전트를 Gateway를 통해 외부 API, 도구 또는 서비스에 연결하거나 Cedar 정책으로 도구 접근을 제한할 때 사용합니다. 게이트웨이 설정, 대상...
aurora-dsql
aws
Aurora DSQL 클러스터를 프로비저닝하고 관리하며, psql 또는 DSQL 커넥터를 통해 연결하고, 스키마를 관리하고, 쿼리를 실행하고, MySQL에서 마이그레이션하고, 쿼리 계획을 진단합니다.
transitgateway
aws
AWS Transit Gateway를 구성합니다: 허브를 생성하고 VPC를 연결하며, 라우팅 테이블로 트래픽을 분리하고, 허브를 통해 이그레스 및 검사를 중앙화합니다…