querying-aws-sagemaker-catalog

작성자: aws

SageMaker Catalog 자산 메타데이터 테이블에서 SQL 분석을 실행하며, S3 Tables에서 Apache Iceberg로 내보낸 데이터를 대상으로 합니다. 거버넌스 쿼리, 자산 성장 추적 등을 다룹니다.

npx skills add https://github.com/aws/agent-toolkit-for-aws --skill querying-aws-sagemaker-catalog

Query AWS SageMaker Catalog System Tables

Overview

Works best with the AWS MCP server for sandboxed execution and audit logging. All commands below use the AWS CLI and work in any environment with configured AWS credentials.

Amazon SageMaker Unified Studio (whose catalog feature is referred to below as SageMaker Catalog) exports asset metadata as a daily-snapshot Apache Iceberg table in the AWS-managed aws-sagemaker-catalog table bucket. This enables SQL queries over your entire data catalog inventory — asset counts, governance gaps, ownership audits, and historical comparisons — without building custom ETL.

Data is partitioned by snapshot_time and exported once daily (around midnight per region). The table is read-only.

Decision Tree

User intentUse this skill?Alternative
SQL analytics on catalog state (counts, governance, trends)Yes
Historical comparison ("what changed in catalog last week")Yes — time travel via snapshot_time
Find assets without owners or descriptionsYes
Find a specific table by name or conceptNofinding-data-lake-assets or Glue Discovery search
Browse/enumerate catalog interactivelyNoexploring-data-catalog
Run a query on a table's dataNoquerying-data-lake
Manage catalog metadata (add descriptions, tags)NoGlue Discovery put-form-type / associate-glossary-terms

Common Tasks

1. Check If Configured

aws datazone get-data-export-configuration \
  --domain-identifier <DOMAIN_ID> \
  --region <REGION>
  • If no domain exists: aws datazone list-domains --region <REGION>
  • If export not enabled: guide user to enable.
  • One domain per account per region.

Verify table bucket exists:

aws s3tables list-table-buckets --region <REGION> \
  --query "tableBuckets[?name=='aws-sagemaker-catalog']"

2. Enable

With KMS encryption (recommended for production):

aws datazone put-data-export-configuration \
  --domain-identifier <DOMAIN_ID> \
  --region <REGION> \
  --enable-export \
  --encryption-configuration kmsKeyArn=<KMS_KEY_ARN>,sseAlgorithm=aws:kms

Note: Encryption cannot be changed after creation. Always specify KMS for sensitive catalog data.

Without encryption (for quick testing only):

aws datazone put-data-export-configuration \
  --domain-identifier <DOMAIN_ID> \
  --region <REGION> \
  --enable-export

First data available within 24 hours. See: Exporting asset metadata

3. Verify Permissions for Querying

Requires:

  • S3 Tables federated catalog registered in Glue (s3tablescatalog)
  • Lake Formation SELECT + DESCRIBE grants on the table

Grant access:

aws lakeformation grant-permissions \
  --principal DataLakePrincipalIdentifier=<ROLE_ARN> \
  --resource '{"Table": {"CatalogId": "<ACCOUNT>:s3tablescatalog/aws-sagemaker-catalog", "DatabaseName": "asset_metadata", "Name": "asset"}}' \
  --permissions DESCRIBE SELECT \
  --region <REGION>

4. Query

Query syntax:

"s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"

Constraints:

  • You MUST always filter by snapshot_time — without it, the query scans all historical snapshots and returns duplicates

  • You MUST confirm workgroup and output location before executing

  • Default to DATE(snapshot_time) = CURRENT_DATE for current state

  • You SHOULD use the key columns documented in this skill to build queries. If you need the full schema, run get-tables once:

    aws glue get-tables --catalog-id "<ACCOUNT>:s3tablescatalog/aws-sagemaker-catalog" --database-name "asset_metadata" --region <REGION>
    

Key columns:

ColumnWhat it holdsUsage
snapshot_timePartition key — daily snapshot timestampAlways filter on this
asset_idUnique catalog asset identifierPrimary key for lookups
resource_type_enumGlueTable, RedshiftTable, S3Collection, etc.Filter by asset type
resource_idARN or native identifierCross-reference with source systems
asset_nameBusiness-friendly nameDisplay, search
resource_nameTechnical name (table name, prefix)Filtering
business_descriptionBusiness context (NULL if not provided)Governance gaps
extended_metadatamap<string,string> — flexible key-value attributesUse bracket notation: extended_metadata['owningEntityId']
asset_created_timeWhen asset first appeared in catalogGrowth analysis
asset_updated_timeLast modification timeFreshness checks

Current catalog state:

SELECT resource_type_enum, COUNT(*) as count
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) = CURRENT_DATE
GROUP BY resource_type_enum
ORDER BY count DESC;

Assets without business descriptions:

SELECT asset_name, resource_name, resource_type_enum, account_id
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) = CURRENT_DATE
  AND business_description IS NULL;

Asset growth over last 30 days:

SELECT DATE(snapshot_time) as date, COUNT(*) as total_assets
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) >= CURRENT_DATE - INTERVAL '30' DAY
GROUP BY DATE(snapshot_time)
ORDER BY date DESC;

Time travel — compare current vs 7 days ago (new descriptions added):

SELECT t.asset_id, t.resource_name,
       p.business_description as before,
       t.business_description as now
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset" t
JOIN "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset" p
  ON t.asset_id = p.asset_id
WHERE DATE(t.snapshot_time) = CURRENT_DATE
  AND DATE(p.snapshot_time) = CURRENT_DATE - INTERVAL '7' DAY
  AND p.business_description IS NULL
  AND t.business_description IS NOT NULL;

Assets by owner:

SELECT extended_metadata['owningEntityId'] as owner, COUNT(*) as count
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) = CURRENT_DATE
  AND extended_metadata['owningEntityId'] IS NOT NULL
GROUP BY extended_metadata['owningEntityId']
ORDER BY count DESC;

Filter by metadata form field:

SELECT *
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) = CURRENT_DATE
  AND extended_metadata['<metadata-form-name>.<field-name>'] = '<field-value>';

Key Behaviors

  • Daily snapshots — exported around midnight per region
  • Always filter by snapshot_time — without it you get all history (duplicates, slow)
  • One domain per account per region — to switch domains, delete config first
  • No additional charge beyond S3 Tables storage + Athena queries
  • Read-only — to update asset metadata, use Glue Discovery APIs or SageMaker Unified Studio

Troubleshooting

ErrorCauseFix
aws-sagemaker-catalog bucket not foundExport not enabledRun put-data-export-configuration --enable-export
Empty results with CURRENT_DATEFirst export hasn't run yet (takes up to 24h)Wait; try yesterday's date
AccessDenied on queryMissing Lake Formation grantsGrant SELECT + DESCRIBE on the table
CATALOG_NOT_FOUNDS3 Tables not registered in GlueEnable integration: S3 console > Table buckets > Enable integration
Duplicate rows in resultsMissing snapshot_time filterAdd WHERE DATE(snapshot_time) = CURRENT_DATE
extended_metadata key returns NULLKey doesn't exist for that assetCheck available keys: SELECT DISTINCT key FROM ... CROSS JOIN UNNEST(map_keys(extended_metadata)) AS t(key) WHERE DATE(snapshot_time) = CURRENT_DATE
Cannot update export encryptionEncryption set at creation time onlyDelete and recreate export config

Security Considerations

Data sensitivity: Catalog metadata exposes organizational structure including asset names, ownership, account IDs, naming conventions, and internal resource identifiers. Treat query results as sensitive by default.

Encryption at rest: Always enable KMS encryption when creating the export configuration. Encryption cannot be changed after creation. Additionally, configure SSE-KMS on your Athena workgroup output bucket.

Least-privilege access: Grant Lake Formation SELECT + DESCRIBE only on the specific asset_metadata.asset table to roles that need catalog analytics. Avoid granting access to the entire aws-sagemaker-catalog bucket.

Audit trail: Enable CloudTrail logging for DataZone (PutDataExportConfiguration, GetDataExportConfiguration), Athena (StartQueryExecution, GetQueryResults), and S3 Tables API calls to track who queries catalog metadata.

Credential hygiene: Use IAM roles with temporary credentials for querying. Avoid long-lived access keys for users accessing catalog metadata. Scope down or rotate principals when access is no longer needed.

Additional Resources

aws의 다른 스킬

agents-build
aws
Use to extend an existing agent project with memory, app integration, VPC, multi-agent, migration, model, browser, code interpreter, payments, or resource…
official
agents-connect
aws
Use when connecting your agent to external APIs, tools, or services via Gateway, or restricting tool access with Cedar policies. Handles gateway setup, target…
official
agents-debug
aws
Use when your agent or environment is broken — wrong answers, errors, timeouts, tool failures, or CLI issues. Reads traces and logs to diagnose root causes.…
official
agents-deploy
aws
에이전트를 AWS에 배포할 때 또는 배포가 실패했을 때 사용합니다. 사전 검증, CDK/IAM/할당량 오류 진단, 버전 관리, 롤백 등을 처리합니다.
official
agents-get-started
aws
개발자가 새 에이전트 프로젝트를 만들거나 AgentCore를 시작하려 할 때 사용합니다. 프레임워크 선택, 프로젝트 스캐폴딩, 첫 배포 등을 처리합니다.
official
agents-harden
aws
Use when preparing your agent for production — IAM scoping, inbound auth (JWT, SigV4), secrets management, cold start optimization, session lifecycle, rate…
official
agents-pay
aws
런타임에 x402로 보호된 콘텐츠에 대해 이 에이전트가 비용을 지불해야 할 때 사용: 작업 중 페이월을 만나면 AgentCore Payments를 통해 결제를 처리하고 적용하는 경우…
official
amazon-aurora-mysql
aws
Amazon Aurora MySQL — Aurora MySQL 클러스터를 구체적으로 생성, 수정, 조언합니다(MySQL 호환 엔진, Aurora serverless, 병렬 쿼리).…
official