querying-data-lake

작성자: aws

기본 및 페더레이션 카탈로그(Glue, S3 Tables, Redshift)에서 Athena SQL 쿼리를 실행하고 관리합니다. "데이터 쿼리", "SQL 실행", "athena" 같은 문구에 반응합니다.

npx skills add https://github.com/aws/agent-toolkit-for-aws --skill querying-data-lake

Query Data Lake

Execute SQL queries on Amazon Athena across default and federated catalogs (Glue, S3 Tables, Redshift) with workgroup selection, statement classification, and error recovery.

Overview

Executes and manages Athena SQL queries across default and federated catalogs. Selects a workgroup, resolves target assets (delegating fuzzy references to finding-data-lake-assets), classifies statements for safety, and reports cost and data scanned. Use the AWS MCP server for sandboxed execution and audit logging; the same AWS CLI commands work directly when the MCP server is not available.

Constraints for parameter acquisition:

  • You MUST accept a single optional argument: SQL text, a named-query name, a workgroup name, a catalog name, or profile TABLE_NAME
  • You MUST accept the argument as direct text or a pointer to a file containing SQL
  • You MUST ask the user for the target AWS region if not already set
  • You MUST confirm the output S3 location before executing any non-trivial query
  • You MUST respect the user's decision to abort at any step

Common Tasks

1. Verify Dependencies

Check for required tools and AWS access before running queries.

Constraints:

  • You MUST verify AWS MCP server tools are available (aws___call_aws) and run queries through them when present; fall back to AWS CLI only if the MCP server is unavailable
  • You MUST NOT fall back to shell or Bash for query execution — results must be captured via the MCP tool or aws athena CLI so output location and cost are tracked
  • You MUST confirm credentials with aws sts get-caller-identity and inform the user about any missing tools

2. Resolve Workgroup

Check caller identity, list workgroups, auto-select the best one (see workgroup-selection.md).

Constraints:

  • You MUST select a workgroup before submitting any query (prevents output-location errors)
  • You MUST present the selected workgroup and its output location to the user
  • You MUST NOT auto-escalate to a different workgroup on failure without user confirmation

3. Resolve the Target Asset

If the user refers to a table by name, by business concept ("our quarterly report", "the sales data"), by S3 path, or by catalog without specifying the table, delegate to finding-data-lake-assets to return the concrete database.table (and catalog if non-default).

Constraints:

  • You MUST NOT attempt to resolve fuzzy asset references with athena list-data-catalogs or by iterating get-tables — those miss federated catalogs and waste tokens
  • You SHOULD skip this step only when the user provides a fully-qualified reference (exact database.table) or raw SQL they want executed as-is
  • You MUST state the resolved asset explicitly before building the query: "Found [table] in [catalog]. Using this for the query."
  • You SHOULD default to the default Glue catalog unless the user mentions "federated", "Redshift", "S3 Tables", or finding-data-lake-assets returns a different catalog

4. Discover Schema

For analytical queries, You SHOULD profile the target table before building the final query. You MUST show sample rows (SELECT ... LIMIT 5) as part of profiling.

5. Build Query

Table addressing depends on catalog type:

  • Default Glue catalog: database.table (omit the catalog prefix for single-catalog queries). In cross-catalog queries, qualify default-catalog tables with "awsdatacatalog".database.table.
  • Registered data source: datasource.database.table
  • Unregistered Glue catalog: "catalog/subcatalog".database.table

6. Classify and Execute

Classify the SQL statement before executing:

StatementBehavior
SELECT, SHOW, DESCRIBE, EXPLAINSafe — execute
INSERT, UPDATE, DELETE, DROP, ALTER, CREATE, TRUNCATE, MERGEDestructive — warn the user and require explicit confirmation
UnsureTreat as destructive; confirm

Example tool call (via AWS MCP server):

aws___call_aws(command="aws athena start-query-execution --work-group <WORKGROUP_NAME> --query-string '<sql>' --query-execution-context Database=<db>")

For federated or S3 Tables catalogs, also set Catalog=<CATALOG_PATH> in the execution context (e.g. Catalog=s3tablescatalog/<BUCKET_NAME>).

Constraints:

  • You MUST warn the user before executing when the target is Redshift-federated ("No partition pruning — every query scans the full table")
  • You MUST warn the user before executing a cross-catalog join ("Cross-catalog joins incur network overhead and may be slow")
  • You MUST confirm the output S3 location before executing
  • You MUST explain which tool is being called before executing
  • You MUST respect the user's decision to abort

7. Present and Recover

Present results with cost, data scanned, duration, and actionable insights. On failure, list available workgroups and let the user choose which to retry with.

Argument Routing

Resolve in this order; stop at the first match:

  1. Contains SQL keywords (SELECT, SHOW, DESCRIBE, INSERT, etc.) — SQL text, execute directly
  2. profile TABLE_NAME — run comprehensive table profiling (see query-patterns.md)
  3. Matches a known named query — look up and execute
  4. Matches a known workgroup — show workgroup status and recent queries
  5. Matches a known catalog — delegate to exploring-data-catalog to enumerate databases and tables
  6. No args — show recent query activity and available tables

Principles

  • Always select workgroup before executing (prevents output-location errors)
  • Profile unfamiliar tables before running analytical queries
  • Present cost alongside results so users build cost awareness
  • Suggest LIMIT for exploratory queries on large tables
  • Never ask domain questions with obvious answers, but always confirm security-relevant actions (workgroup switches, output location changes, non-SELECT statements)

Troubleshooting

ErrorCauseFix
Redshift identifier error with mixed caseRedshift-federated names are lowercase onlyLowercase the identifier
CatalogId validation failureARN passed instead of catalog namePass the catalog name, not the ARN
Cross-catalog information_schema returns nothingMissing catalog qualifierUse catalog-qualified path: "catalog".information_schema.tables
Query fails with output-location errorWorkgroup has no output location configuredSelect a different workgroup with an output location, or configure one
Destructive statement executed without confirmationStatement classification skippedAlways classify INSERT/UPDATE/DELETE/DROP/ALTER/CREATE/TRUNCATE/MERGE and confirm with the user

Additional Resources

aws의 다른 스킬

analyzing-release-readiness
aws
GitHub PR, GitLab MR 또는 로컬 브랜치에서 병합 전 릴리스 준비 검토를 트리거합니다. 사용자가 코드 변경 사항의 위험성, 정확성 등을 분석하려 할 때 사용합니다.
scanning-with-aws-security-agent
aws
작업 공간에서 AWS Security Agent 스캔 실행 — 소스를 AWS에 업로드하고, 관리형 Security Agent 서비스로 스캔한 후, 순위가 매겨진 검증된 결과를 반환합니다…
coordinating-multi-space-devops-agent
aws
하나의 Claude Code 세션에서 여러 AgentSpaces에 걸쳐 AWS DevOps Agent를 조정하세요 — 질문을 올바른 공간(프로덕션 vs 스테이징 vs 지식)으로 라우팅하고,…
aws-security
aws
AWS 보안 서비스 및 워크플로우를 다룹니다 — Security Hub V2 (OCSF) findings, 커넥터, 애그리게이터, 자동화 규칙, 보안 상태 요약 등…
querying-aws-sagemaker-catalog
aws
SageMaker Catalog 자산 메타데이터 테이블에서 SQL 분석을 실행하며, S3 Tables에서 Apache Iceberg로 내보낸 데이터를 대상으로 합니다. 거버넌스 쿼리, 자산 성장 추적 등을 다룹니다.
agents-connect
aws
에이전트를 Gateway를 통해 외부 API, 도구 또는 서비스에 연결하거나 Cedar 정책으로 도구 접근을 제한할 때 사용합니다. 게이트웨이 설정, 대상...
aurora-dsql
aws
Aurora DSQL 클러스터를 프로비저닝하고 관리하며, psql 또는 DSQL 커넥터를 통해 연결하고, 스키마를 관리하고, 쿼리를 실행하고, MySQL에서 마이그레이션하고, 쿼리 계획을 진단합니다.
transitgateway
aws
AWS Transit Gateway를 구성합니다: 허브를 생성하고 VPC를 연결하며, 라우팅 테이블로 트래픽을 분리하고, 허브를 통해 이그레스 및 검사를 중앙화합니다…