storing-and-querying-vectors

작성자: aws

Amazon S3 Vectors를 사용하여 벡터 임베딩을 저장하고 쿼리합니다. 자체 API 네임스페이스(s3vectors)를 갖춘 비용 효율적인 장기 벡터 스토리지 서비스입니다. 트리거…

npx skills add https://github.com/aws/agent-toolkit-for-aws --skill storing-and-querying-vectors

Store and Query Vectors with Amazon S3 Vectors

Overview

Amazon S3 Vectors is a cost-effective AWS service for storing and querying vector embeddings at scale. Optimized for long-term storage with subsecond latency for cold queries, as low as 100ms for warm queries.

Decision Guide

  • Hundreds/thousands of sustained queries per second (QPS): Wrong tool. Recommend OpenSearch.
  • Hybrid search, aggregations, faceted search: Recommend OpenSearch with S3 Vectors as storage engine. For OpenSearch integration, search AWS docs for "Using S3 Vectors with OpenSearch Service".
  • Tiered (bulk + hot): S3 Vectors for storage + OpenSearch Serverless for real-time. See references/limits-and-patterns.md.
  • Cost-effective storage, infrequent queries, RAG: S3 Vectors is the right fit. Proceed.

For latest guidance, search AWS docs for "S3 Vectors best practices".

Common Tasks

Classify the request before starting:

  • Simple query: Existing index, skip to Step 6
  • Standard: You MUST list existing indexes first and suggest reusing if relevant. Else, new index + store vectors, follow Steps 2-6
  • Migration or multi-tenant: Read references/limits-and-patterns.md first, then Steps 2-6

You MUST execute commands using AWS MCP server tools when connected. Fall back to AWS CLI only if AWS MCP is unavailable. You MUST explain each step to the user before executing.

1. Verify Dependencies

Constraints:

  • You MUST check whether AWS MCP tools or AWS CLI is available and inform user if missing
  • You MUST confirm target AWS region

2. Create a Vector Bucket

You MUST confirm bucket name with user. Names: 3-63 chars, lowercase letters, numbers, hyphens only. Encryption (SSE-S3 default or SSE-KMS for compliance) is immutable after creation.

aws s3vectors create-vector-bucket \
  --vector-bucket-name <BUCKET_NAME>

Constraints:

  • You MUST explain encryption cannot be changed after creation
  • For SSE-KMS, KMS key policy MUST grant kms:GenerateDataKey and kms:Decrypt to the S3 Vectors service principal indexing.s3vectors.amazonaws.com. You MUST use full KMS key ARN (not alias). See references/limits-and-patterns.md for command example.

3. Create a Vector Index

Every parameter is immutable after creation.

Pre-flight checklist (confirm ALL with user):

  1. Dimension (required, integer 1-4096) -- MUST match embedding model output
  2. Distance metric (required) -- cosine or euclidean. Use embedding model's recommended metric;
  3. Non-filterable metadata keys (optional, max 10, 1-63 chars) -- Declare at creation or lose forever. For Bedrock Knowledge Bases integration, search AWS docs for "S3 Vectors Bedrock Knowledge Bases prerequisites" to get the required key names.
  4. Encryption (optional) -- Inherits from bucket. Override per-index if needed.
aws s3vectors create-index \
  --vector-bucket-name <BUCKET_NAME> \
  --index-name <INDEX_NAME> \
  --dimension <DIM> \
  --distance-metric <cosine|euclidean> \
  --data-type float32 \
  --metadata-configuration '{"nonFilterableMetadataKeys":["<KEY1>","<KEY2>"]}'

Omit --metadata-configuration if no non-filterable keys are needed.

Index names: 3-63 chars, lowercase, numbers, hyphens, dots. Unique within bucket. Filterable metadata: 2 KB limit. Total metadata (filterable + non-filterable combined): 40 KB. See references/metadata-filtering.md.

4. Generate Embeddings (if needed)

Skip to Step 5 (store) or Step 6 (query) if user already has embeddings.

Constraints:

  • You MUST ask which embedding model to use if not specified
  • You MUST NOT assume a default model
  • Dimension MUST match Step 3
  • You MUST use the same model for both storing and querying

Generate embeddings with Bedrock invoke-model:

aws bedrock-runtime invoke-model \
  --model-id <MODEL_ID> \
  --content-type application/json \
  --cli-binary-format raw-in-base64-out \
  --body '{"inputText": "your text"}' \
  invoke-model-output.json

You MUST use --cli-binary-format raw-in-base64-out for CLI v2. Output file is required for CLI. The response key is model-dependent (e.g., embedding for Titan, embeddings for Cohere). For Titan, parse with json.load(open('invoke-model-output.json'))['embedding']. Use embedding array as float32 in put-vectors or query-vectors. For batch embedding generation, use AWS SDK or CLI.

5. Put Vectors

aws s3vectors put-vectors \
  --vector-bucket-name <BUCKET_NAME> \
  --index-name <INDEX_NAME> \
  --vectors '[{"key":"<ID>","data":{"float32":[<EMBEDDING>]},"metadata":{"topic":"science"}}]'

Constraints:

  • You MUST NOT exceed 500 vectors per call
  • You SHOULD batch vectors for cost optimization
  • For bulk operations, You SHOULD use an SDK instead of CLI -- vector payloads may be too large for shell arguments
  • You MUST implement retry with backoff on 429 TooManyRequestsException
  • See references/limits-and-patterns.md for batch patterns

6. Query Vectors

Generate embedding if needed (Step 4), then query:

aws s3vectors query-vectors \
  --vector-bucket-name <BUCKET_NAME> \
  --index-name <INDEX_NAME> \
  --query-vector '{"float32":[<EMBEDDING>]}' \
  --top-k 10 \
  --return-distance

Optional: add --return-metadata and/or --filter '{"topic":{"$eq":"science"}}' (both require GetVectors permission). See references/metadata-filtering.md.

Example response body: {"vectors": [{"key": "id1", "distance": 0.45, "metadata": {"topic": "science"}}, ...], "distanceMetric": "cosine"}

Constraints:

  • Using --filter or --return-metadata requires both s3vectors:QueryVectors AND s3vectors:GetVectors IAM permissions. Without GetVectors, these options return 403.

Troubleshooting

ErrorCauseFix
DimensionMismatchDims don't match indexUse matching model, or delete/recreate index (confirm with user -- destroys all vectors).
403 Forbidden with --filter or --return-metadataMissing s3vectors:GetVectorsAdd s3vectors:GetVectors to IAM policy.
Fewer results than --top-kFew vectors match filterExpected -- filtering is inline. Broaden filter.
429 TooManyRequestsExceptionExceeded per-index rate limitsRetry with backoff. Shard across indexes for sustained throughput. Search AWS docs for "S3 Vectors limitations and restrictions" for current limits.
AccessDeniedExceptionMissing s3vectors:* IAM actionsS3 Vectors uses s3vectors:* namespace, not s3:*. Update IAM policy.
RequestTimeoutException or service unavailableRequest timeout or region not supportedRetry request. For regional availability, search AWS docs for "S3 Vectors limitations and restrictions".

Additional Resources

aws의 다른 스킬

analyzing-release-readiness
aws
GitHub PR, GitLab MR 또는 로컬 브랜치에서 병합 전 릴리스 준비 검토를 트리거합니다. 사용자가 코드 변경 사항의 위험성, 정확성 등을 분석하려 할 때 사용합니다.
scanning-with-aws-security-agent
aws
작업 공간에서 AWS Security Agent 스캔 실행 — 소스를 AWS에 업로드하고, 관리형 Security Agent 서비스로 스캔한 후, 순위가 매겨진 검증된 결과를 반환합니다…
coordinating-multi-space-devops-agent
aws
하나의 Claude Code 세션에서 여러 AgentSpaces에 걸쳐 AWS DevOps Agent를 조정하세요 — 질문을 올바른 공간(프로덕션 vs 스테이징 vs 지식)으로 라우팅하고,…
aws-security
aws
AWS 보안 서비스 및 워크플로우를 다룹니다 — Security Hub V2 (OCSF) findings, 커넥터, 애그리게이터, 자동화 규칙, 보안 상태 요약 등…
querying-aws-sagemaker-catalog
aws
SageMaker Catalog 자산 메타데이터 테이블에서 SQL 분석을 실행하며, S3 Tables에서 Apache Iceberg로 내보낸 데이터를 대상으로 합니다. 거버넌스 쿼리, 자산 성장 추적 등을 다룹니다.
agents-connect
aws
에이전트를 Gateway를 통해 외부 API, 도구 또는 서비스에 연결하거나 Cedar 정책으로 도구 접근을 제한할 때 사용합니다. 게이트웨이 설정, 대상...
aurora-dsql
aws
Aurora DSQL 클러스터를 프로비저닝하고 관리하며, psql 또는 DSQL 커넥터를 통해 연결하고, 스키마를 관리하고, 쿼리를 실행하고, MySQL에서 마이그레이션하고, 쿼리 계획을 진단합니다.
transitgateway
aws
AWS Transit Gateway를 구성합니다: 허브를 생성하고 VPC를 연결하며, 라우팅 테이블로 트래픽을 분리하고, 허브를 통해 이그레스 및 검사를 중앙화합니다…