redshift-guide

작성자: aws

Amazon Redshift는 PostgreSQL이 아닙니다 — PostgreSQL에서 파생된 LLM 실수를 교정합니다; Redshift 고유의 SQL, DDL, COPY/UNLOAD, 시스템 뷰, 메타데이터 탐색 등을 다룹니다…

npx skills add https://github.com/aws/agent-toolkit-for-aws --skill redshift-guide

Amazon Redshift Guide

Redshift is NOT PostgreSQL (read first)

Redshift speaks PostgreSQL's wire protocol and shares much of its surface syntax, so LLMs assume PostgreSQL behavior carries over — it frequently does not. Divergences span system tables (pg_catalog is incomplete), DDL (no indexes, no sequences), functions (string_agg, SUBSTR on tables, leader-node-only functions), types (a text column becomes VARCHAR(256)), and comparison semantics (trailing blanks, unenforced constraints). Assume divergence and verify against the reference below — do not answer from PostgreSQL habit. Common PostgreSQL→Redshift divergences are in references/redshift-sql-syntax.md.

Works best with the AWS MCP server — it runs the AWS CLI and Redshift Data API calls below in a sandboxed, audit-logged environment. All guidance here is plain AWS CLI and SQL and works without it.

STEP 0: Serverless or Provisioned?

Establish this before answering — APIs, system tables, and capabilities differ. Take it from the question when it says which one; ask when it does not. SELECT version() does not identify it.

  • Serverless — identified by a workgroup (and namespace). Data API calls take --workgroup-name; the user says "workgroup"/"Serverless".
  • Provisioned — identified by a cluster. Data API calls take --cluster-identifier; the user says "cluster".
TargetSystem ViewsCredentials API
ProvisionedSYS_, all SVV_ + STL_, STV_, SVL_, SVCS_ (single-AZ only — disabled on Multi-AZ)redshift:GetClusterCredentials
ServerlessSYS_ + a subset of SVV_ ONLY (no STL/STV/SVL/SVCS)redshift-serverless:GetCredentials

Critical Facts

  • SHOW commands are the primary metadata interface — SHOW DATABASES, SHOW SCHEMAS, SHOW TABLES, SHOW COLUMNS, SHOW TABLE, SHOW VIEW. Do NOT default to pg_catalog or information_schema. → Load references/redshift-sql-metadata.md for metadata/discovery questions and any "relation does not exist" report — it has the diagnostic flow.
  • SYS_ views are the preferred system views — they work everywhere. STL_, STV_, SVL_, and SVCS_ are provisioned single-AZ only, and some SVV_ views are unsupported on Serverless. → Load references/redshift-sql-metadata.md for any system-view or monitoring question.
  • sys_load_error_detail for COPY debugging (not stl_load_errors, which is provisioned single-AZ only).
  • DATEADD/DATEDIFF — unit-first argument order: DATEADD(day, -30, GETDATE()), DATEDIFF(day, start, end).
  • APPROXIMATE COUNT(DISTINCT col) — Redshift-specific, ~2% error, much faster than exact COUNT(DISTINCT) on large datasets.
  • MERGE ... REMOVE DUPLICATES — simplified dedup when source and target have identical schemas.
  • COPY should use IAM_ROLE (the namespace role, not the caller role) + supports MANIFEST for explicit file lists + MAXERROR for error tolerance.
  • SUBSTR() is leader-node-only — works on literals but errors on table columns (SUBSTR() function is not supported (Hint: use SUBSTRING instead)). Use SUBSTRING() on columns.
  • UNIQUE / PRIMARY KEY / FOREIGN KEY are informational only — NOT enforced (duplicate rows are accepted with no error). Optimizer hints; enforce integrity in the application or via MERGE. NOT NULL IS enforced.
  • SHOW VIEW <schema.name> returns the definition of a regular view, materialized view, or late-binding view. MV freshness: SVV_MV_INFO (is_stale).
  • TOP N and LIMIT N both work (TOP N PERCENT does not). A text column becomes VARCHAR(256) — use VARCHAR(max) or explicit length.
  • Iceberg tables use CREATE TABLE ... USING ICEBERG (not STORED AS ICEBERG, not TABLE_FORMAT=ICEBERG).
  • Datashares support read and write operations — consumers can write once the producer grants write privileges. Treat "permission denied" on a datashare write as a missing grant, not an unsupported operation. → Load references/redshift-sql-metadata.md for requirements and limits.

Safety Guardrails

BLOCK: DROP DATABASE, DELETE without WHERE, publicly-accessible=true, GRANT ALL ON ALL WARN then confirm: RESIZE, RESTORE, VACUUM on large tables, ALTER PASSWORD, WLM config change Confirm: CREATE, GRANT specific, COPY, UNLOAD

Security Considerations

Apply these defaults when generating anything that connects, loads, or exports. Details are in the reference files noted.

  • In transit: the Data API is HTTPS-only. For JDBC/ODBC set the require_ssl parameter and connect with sslmode=verify-full so the server certificate is checked.
  • At rest: keep cluster/namespace encryption enabled, and add ENCRYPTED KMS_KEY_ID '<arn>' to UNLOAD — it writes query results to S3, outside Redshift's own encryption. → references/redshift-sql-ddl-copy.md
  • Credentials: prefer SecretArn (Secrets Manager) or IAM Identity Center; DbUser is acceptable because it issues temporary credentials. Never place database passwords in code, environment variables, or SQL text. → references/redshift-sql-recipes-load-api.md
  • Least privilege: scope the namespace IAM_ROLE to the specific bucket and prefix (s3:GetObject on arn:aws:s3:::<bucket>/<prefix>/*), not s3:* or a managed full-access policy, and condition its trust policy on both aws:SourceArn (the cluster/namespace ARN) and aws:SourceAccount — SourceArn alone still allows another resource in the account to assume it. Grant per-object privileges rather than GRANT ALL ON ALL.
  • Audit: CloudTrail records redshift-data:* API calls but not the SQL executed; enable Redshift audit logging (useractivitylog, connectionlog, userlog) for that. Both capture query text and user activity, so encrypt every destination in use: the CloudWatch Logs group (aws logs associate-kms-key), the CloudTrail trail (SSE-KMS), and the audit-log S3 bucket (SSE-S3 — audit logging to S3 supports only S3-managed keys, not KMS). Serverless only supports sending audit logs to CloudWatch.
  • Network: keep PubliclyAccessible=false and connect over a VPC endpoint. Do not open port 5439 to 0.0.0.0/0 or ::/0 — scope inbound rules to specific CIDRs or to a referencing security group.
  • Sensitive data: Data API results persist for 24h and sys_load_error_detail can echo fragments of rejected rows, so treat statement IDs and load-error output as sensitive.
  • Further reading: Security in Amazon Redshift for the full guidance behind these defaults.

Routing Table

MANDATORY: When a question matches a row below, you MUST load and read the referenced file BEFORE answering.

Ask whether the target is provisioned or Serverless before giving troubleshooting steps — unless the question already says which one, in which case use that and do not re-confirm.

User IntentRoute To
"CREATE TABLE", "DISTKEY/SORTKEY", "ENCODE", "IDENTITY", "COPY", "UNLOAD", "IAM_ROLE", "Iceberg table"references/redshift-sql-ddl-copy.md
"LISTAGG", "DATEADD/DATEDIFF", "NVL/DECODE", "type mapping", "text type", "VARBYTE", "recursive CTE"references/redshift-sql-functions-types.md
"QUALIFY", "PIVOT/UNPIVOT", "MERGE", "TOP N", "SUBSTR error", "UNIQUE/PK not enforced", "trailing blanks", "leader-node function", "JSON", "SUPER", "PartiQL", "nested/semi-structured data"references/redshift-sql-extensions-semantics.md
"system view", "SVV_/SYS_", "SHOW commands", "STL vs SYS", "list tables", "distkey/sortkey lookup", "datashare discovery", "2-part vs 3-part", "permission denied", "GRANT", "privileges", "relation/table does not exist"references/redshift-sql-metadata.md
"how do I write SQL", "PostgreSQL vs Redshift", "which SQL reference", general dialect questionreferences/redshift-sql-syntax.md (index of the 6 SQL references + PostgreSQL-vs-Redshift failure table)
"COPY failed", "load error", "Data API poll", "async query", "Data API throttle"references/redshift-sql-recipes-load-api.md
"materialized view", "MV refresh", "AUTO REFRESH", "stale view"references/redshift-sql-materialized-views.md
General Redshift question not matching aboveAnswer directly from general knowledge
Aurora, RDS, DynamoDB, Athena (non-Redshift)REFUSE. State this skill is for Amazon Redshift only. Do not provide guidance for other database services.

Data API Quick Reference

→ Load references/redshift-sql-recipes-load-api.md before answering ANY Data API, COPY-error, or async-query question. It carries the bounded poll loop, the HasResultSet and ResourceNotFoundException handling, the per-target parameters, and the auth options.

Data API calls are async by default — use long polling (--wait-time-seconds, 1–30) rather than blind sleeps, and keep a bounded loop for work that can exceed 30s. Serverless takes --workgroup-name, provisioned takes --cluster-identifier.

aws의 다른 스킬

analyzing-release-readiness
aws
GitHub PR, GitLab MR 또는 로컬 브랜치에서 병합 전 릴리스 준비 검토를 트리거합니다. 사용자가 코드 변경 사항의 위험성, 정확성 등을 분석하려 할 때 사용합니다.
scanning-with-aws-security-agent
aws
작업 공간에서 AWS Security Agent 스캔 실행 — 소스를 AWS에 업로드하고, 관리형 Security Agent 서비스로 스캔한 후, 순위가 매겨진 검증된 결과를 반환합니다…
coordinating-multi-space-devops-agent
aws
하나의 Claude Code 세션에서 여러 AgentSpaces에 걸쳐 AWS DevOps Agent를 조정하세요 — 질문을 올바른 공간(프로덕션 vs 스테이징 vs 지식)으로 라우팅하고,…
aws-security
aws
AWS 보안 서비스 및 워크플로우를 다룹니다 — Security Hub V2 (OCSF) findings, 커넥터, 애그리게이터, 자동화 규칙, 보안 상태 요약 등…
querying-aws-sagemaker-catalog
aws
SageMaker Catalog 자산 메타데이터 테이블에서 SQL 분석을 실행하며, S3 Tables에서 Apache Iceberg로 내보낸 데이터를 대상으로 합니다. 거버넌스 쿼리, 자산 성장 추적 등을 다룹니다.
agents-connect
aws
에이전트를 Gateway를 통해 외부 API, 도구 또는 서비스에 연결하거나 Cedar 정책으로 도구 접근을 제한할 때 사용합니다. 게이트웨이 설정, 대상...
aurora-dsql
aws
Aurora DSQL 클러스터를 프로비저닝하고 관리하며, psql 또는 DSQL 커넥터를 통해 연결하고, 스키마를 관리하고, 쿼리를 실행하고, MySQL에서 마이그레이션하고, 쿼리 계획을 진단합니다.
transitgateway
aws
AWS Transit Gateway를 구성합니다: 허브를 생성하고 VPC를 연결하며, 라우팅 테이블로 트래픽을 분리하고, 허브를 통해 이그레스 및 검사를 중앙화합니다…