querying-aws-sagemaker-catalog

por aws

Executa análises SQL em tabelas de metadados do catálogo SageMaker Catalog exportadas como Apache Iceberg em S3 Tables. Abrange consultas de governança, acompanhamento de crescimento de ativos,…

npx skills add https://github.com/aws/agent-toolkit-for-aws --skill querying-aws-sagemaker-catalog

Query AWS SageMaker Catalog System Tables

Overview

Works best with the AWS MCP server for sandboxed execution and audit logging. All commands below use the AWS CLI and work in any environment with configured AWS credentials.

Amazon SageMaker Unified Studio (whose catalog feature is referred to below as SageMaker Catalog) exports asset metadata as a daily-snapshot Apache Iceberg table in the AWS-managed aws-sagemaker-catalog table bucket. This enables SQL queries over your entire data catalog inventory — asset counts, governance gaps, ownership audits, and historical comparisons — without building custom ETL.

Data is partitioned by snapshot_time and exported once daily (around midnight per region). The table is read-only.

Decision Tree

User intentUse this skill?Alternative
SQL analytics on catalog state (counts, governance, trends)Yes
Historical comparison ("what changed in catalog last week")Yes — time travel via snapshot_time
Find assets without owners or descriptionsYes
Find a specific table by name or conceptNofinding-data-lake-assets or Glue Discovery search
Browse/enumerate catalog interactivelyNoexploring-data-catalog
Run a query on a table's dataNoquerying-data-lake
Manage catalog metadata (add descriptions, tags)NoGlue Discovery put-form-type / associate-glossary-terms

Common Tasks

1. Check If Configured

aws datazone get-data-export-configuration \
  --domain-identifier <DOMAIN_ID> \
  --region <REGION>
  • If no domain exists: aws datazone list-domains --region <REGION>
  • If export not enabled: guide user to enable.
  • One domain per account per region.

Verify table bucket exists:

aws s3tables list-table-buckets --region <REGION> \
  --query "tableBuckets[?name=='aws-sagemaker-catalog']"

2. Enable

With KMS encryption (recommended for production):

aws datazone put-data-export-configuration \
  --domain-identifier <DOMAIN_ID> \
  --region <REGION> \
  --enable-export \
  --encryption-configuration kmsKeyArn=<KMS_KEY_ARN>,sseAlgorithm=aws:kms

Note: Encryption cannot be changed after creation. Always specify KMS for sensitive catalog data.

Without encryption (for quick testing only):

aws datazone put-data-export-configuration \
  --domain-identifier <DOMAIN_ID> \
  --region <REGION> \
  --enable-export

First data available within 24 hours. See: Exporting asset metadata

3. Verify Permissions for Querying

Requires:

  • S3 Tables federated catalog registered in Glue (s3tablescatalog)
  • Lake Formation SELECT + DESCRIBE grants on the table

Grant access:

aws lakeformation grant-permissions \
  --principal DataLakePrincipalIdentifier=<ROLE_ARN> \
  --resource '{"Table": {"CatalogId": "<ACCOUNT>:s3tablescatalog/aws-sagemaker-catalog", "DatabaseName": "asset_metadata", "Name": "asset"}}' \
  --permissions DESCRIBE SELECT \
  --region <REGION>

4. Query

Query syntax:

"s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"

Constraints:

  • You MUST always filter by snapshot_time — without it, the query scans all historical snapshots and returns duplicates

  • You MUST confirm workgroup and output location before executing

  • Default to DATE(snapshot_time) = CURRENT_DATE for current state

  • You SHOULD use the key columns documented in this skill to build queries. If you need the full schema, run get-tables once:

    aws glue get-tables --catalog-id "<ACCOUNT>:s3tablescatalog/aws-sagemaker-catalog" --database-name "asset_metadata" --region <REGION>
    

Key columns:

ColumnWhat it holdsUsage
snapshot_timePartition key — daily snapshot timestampAlways filter on this
asset_idUnique catalog asset identifierPrimary key for lookups
resource_type_enumGlueTable, RedshiftTable, S3Collection, etc.Filter by asset type
resource_idARN or native identifierCross-reference with source systems
asset_nameBusiness-friendly nameDisplay, search
resource_nameTechnical name (table name, prefix)Filtering
business_descriptionBusiness context (NULL if not provided)Governance gaps
extended_metadatamap<string,string> — flexible key-value attributesUse bracket notation: extended_metadata['owningEntityId']
asset_created_timeWhen asset first appeared in catalogGrowth analysis
asset_updated_timeLast modification timeFreshness checks

Current catalog state:

SELECT resource_type_enum, COUNT(*) as count
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) = CURRENT_DATE
GROUP BY resource_type_enum
ORDER BY count DESC;

Assets without business descriptions:

SELECT asset_name, resource_name, resource_type_enum, account_id
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) = CURRENT_DATE
  AND business_description IS NULL;

Asset growth over last 30 days:

SELECT DATE(snapshot_time) as date, COUNT(*) as total_assets
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) >= CURRENT_DATE - INTERVAL '30' DAY
GROUP BY DATE(snapshot_time)
ORDER BY date DESC;

Time travel — compare current vs 7 days ago (new descriptions added):

SELECT t.asset_id, t.resource_name,
       p.business_description as before,
       t.business_description as now
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset" t
JOIN "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset" p
  ON t.asset_id = p.asset_id
WHERE DATE(t.snapshot_time) = CURRENT_DATE
  AND DATE(p.snapshot_time) = CURRENT_DATE - INTERVAL '7' DAY
  AND p.business_description IS NULL
  AND t.business_description IS NOT NULL;

Assets by owner:

SELECT extended_metadata['owningEntityId'] as owner, COUNT(*) as count
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) = CURRENT_DATE
  AND extended_metadata['owningEntityId'] IS NOT NULL
GROUP BY extended_metadata['owningEntityId']
ORDER BY count DESC;

Filter by metadata form field:

SELECT *
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) = CURRENT_DATE
  AND extended_metadata['<metadata-form-name>.<field-name>'] = '<field-value>';

Key Behaviors

  • Daily snapshots — exported around midnight per region
  • Always filter by snapshot_time — without it you get all history (duplicates, slow)
  • One domain per account per region — to switch domains, delete config first
  • No additional charge beyond S3 Tables storage + Athena queries
  • Read-only — to update asset metadata, use Glue Discovery APIs or SageMaker Unified Studio

Troubleshooting

ErrorCauseFix
aws-sagemaker-catalog bucket not foundExport not enabledRun put-data-export-configuration --enable-export
Empty results with CURRENT_DATEFirst export hasn't run yet (takes up to 24h)Wait; try yesterday's date
AccessDenied on queryMissing Lake Formation grantsGrant SELECT + DESCRIBE on the table
CATALOG_NOT_FOUNDS3 Tables not registered in GlueEnable integration: S3 console > Table buckets > Enable integration
Duplicate rows in resultsMissing snapshot_time filterAdd WHERE DATE(snapshot_time) = CURRENT_DATE
extended_metadata key returns NULLKey doesn't exist for that assetCheck available keys: SELECT DISTINCT key FROM ... CROSS JOIN UNNEST(map_keys(extended_metadata)) AS t(key) WHERE DATE(snapshot_time) = CURRENT_DATE
Cannot update export encryptionEncryption set at creation time onlyDelete and recreate export config

Security Considerations

Data sensitivity: Catalog metadata exposes organizational structure including asset names, ownership, account IDs, naming conventions, and internal resource identifiers. Treat query results as sensitive by default.

Encryption at rest: Always enable KMS encryption when creating the export configuration. Encryption cannot be changed after creation. Additionally, configure SSE-KMS on your Athena workgroup output bucket.

Least-privilege access: Grant Lake Formation SELECT + DESCRIBE only on the specific asset_metadata.asset table to roles that need catalog analytics. Avoid granting access to the entire aws-sagemaker-catalog bucket.

Audit trail: Enable CloudTrail logging for DataZone (PutDataExportConfiguration, GetDataExportConfiguration), Athena (StartQueryExecution, GetQueryResults), and S3 Tables API calls to track who queries catalog metadata.

Credential hygiene: Use IAM roles with temporary credentials for querying. Avoid long-lived access keys for users accessing catalog metadata. Scope down or rotate principals when access is no longer needed.

Additional Resources

Mais skills de aws

agents-build
aws
Use para estender um projeto de agente existente com memória, integração de aplicativos, VPC, multi-agente, migração, modelo, navegador, interpretador de código, pagamentos ou recurso…
official
agents-connect
aws
Use ao conectar seu agente a APIs, ferramentas ou serviços externos via Gateway, ou ao restringir o acesso a ferramentas com políticas Cedar. Lida com a configuração do gateway, destino…
official
agents-debug
aws
Use quando seu agente ou ambiente estiver quebrado — respostas erradas, erros, timeouts, falhas de ferramentas ou problemas de CLI. Lê traces e logs para diagnosticar causas raiz.…
official
agents-deploy
aws
Use ao implantar seu agente na AWS, ou quando uma implantação falhar. Lida com validação de pré-verificação, diagnóstico de erros de CDK/IAM/cota, gerenciamento de versões, rollback,...
official
agents-get-started
aws
Use quando um desenvolvedor quiser criar um novo projeto de agente ou começar com o AgentCore. Lida com seleção de framework, scaffolding de projeto, primeiro deploy e…
official
agents-harden
aws
Use ao preparar seu agente para produção — escopo de IAM, autenticação de entrada (JWT, SigV4), gerenciamento de segredos, otimização de cold start, ciclo de vida de sessão, taxa…
official
agents-pay
aws
Use quando ESTE agente precisar pagar por conteúdo protegido por x402 em tempo de execução: encontrar um paywall no meio da tarefa, resolver isso via AgentCore Payments e aplicar…
official
amazon-aurora-mysql
aws
Amazon Aurora MySQL — cria, modifica e orienta sobre clusters Aurora MySQL especificamente (mecanismo compatível com MySQL, Aurora serverless, consulta paralela).…
official