redis-semantic-cache

작성자: redis

Redis LangCache를 사용하여 Redis Cloud에서 LLM 응답의 시맨틱 캐싱을 수행하는 방법 — SDK 또는 REST API를 통한 검색/설정 호출, 유사도 임계값 조정,…

npx skills add https://github.com/redis/agent-skills --skill redis-semantic-cache

Redis Semantic Cache

Semantic caching for LLM responses with Redis Cloud's LangCache service. Stores prompts as embeddings; subsequent semantically-similar prompts return the cached response without re-calling the model.

LangCache is currently in preview on Redis Cloud. Features and behavior may change.

When to apply

  • Wrapping an LLM call (OpenAI, Anthropic, etc.) with a cache layer to cut cost and latency.
  • Caching RAG answers, classification outputs, or any deterministic LLM workload.
  • Tuning the precision/hit-rate trade-off for a semantic cache.
  • Splitting one application's LLM workloads across multiple cache instances.

1. The cache-aside flow

LangCache fits in front of any LLM call as a standard cache-aside pattern:

  1. Send the user's prompt to LangCache's search.
  2. Cache hit — return the stored response directly.
  3. Cache miss — call the LLM, then set the response so future similar prompts hit.
from langcache import LangCache
import os

lang_cache = LangCache(
    server_url=f"https://{os.getenv('HOST')}",
    cache_id=os.getenv("CACHE_ID"),
    api_key=os.getenv("API_KEY"),
)

result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.9)
if result:
    response = result[0]["response"]
else:
    response = llm.generate("What is Redis?")
    lang_cache.set(prompt="What is Redis?", response=response)

The same operations are available via REST (POST /v1/caches/{cacheId}/entries/search and POST /v1/caches/{cacheId}/entries) when an SDK isn't an option.

See references/langcache-usage.md for full SDK + REST samples and attribute-based storage.

2. Tune the similarity threshold

The threshold controls how close (in embedding cosine distance) a new prompt must be to a cached one to count as a hit. Higher = stricter match, fewer false positives. Lower = more hits, more risk of returning an off-topic answer.

ThresholdBehaviorUse when
0.95+Near-exact match requiredCustomer-facing answers where wrong responses are costly
0.9Balanced defaultMost workloads — start here
0.8Loose semantic matchInternal tools, exploratory queries, FAQ deduplication
# Stricter — fewer false positives
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.95)

# Looser — higher hit rate
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.8)

Adjust by watching the actual cache-hit rate and spot-checking that returned answers are still relevant.

See references/best-practices.md.

3. Separate caches per task type

Different LLM workloads should not share one cache — a "code question" prompt is semantically close to other code questions but has nothing to do with a password-reset support query, and crossing them returns garbage.

support_cache = LangCache(server_url=..., cache_id="support-cache-id", api_key=...)
code_cache    = LangCache(server_url=..., cache_id="code-cache-id",    api_key=...)

Create distinct cache IDs in Redis Cloud per task, and route each call to the right one. As a finer-grained alternative, store and search with custom attributes (e.g. {"category": "database"}) to keep tasks in the same cache but isolated by attribute filter — useful when the same prompt format spans subtopics.

References

redis의 다른 스킬

docs-sync
redis
마스터 브랜치의 구현 및 구성을 분석하여 docs/, README.md, 패키지별 README에서 누락되거나, 부정확하거나, 오래된 문서를 찾습니다.
redis-query-engine
redis
Redis Query Engine (RQE) 가이드: FT.CREATE 스키마 설계, 필드 유형 선택(TEXT, TAG, NUMERIC, GEO, GEOSHAPE, VECTOR), DIALECT 2 쿼리 구문,…
redis-search
redis
Redis Search 가이드: FT.CREATE 스키마 설계, 필드 유형 선택(TEXT, TAG, NUMERIC, GEO, GEOSHAPE, VECTOR, JSON 경로), DIALECT 2 쿼리 구문,…
redis-security
redis
Redis 보안 가이드로 인증(requirepass 및 ACL 사용자), TLS, ACL 기반 최소 권한 접근 제어, 네트워크 노출 제한 등을 다룹니다.
redis-vector-search
redis
Redis 벡터 검색 가이드: HNSW vs FLAT 알고리즘 선택, 벡터 인덱스 구성(차원, 거리 메트릭, 데이터 타입), 필터링된 하이브리드 검색…
bump-test-image
redis
공유 DEFAULT_DOCKER_CONFIG 및 CI 매트릭스에서 기본 Redis docker 테스트 이미지(redislabs/client-libs-test)를 범프한 다음 강제 푸시합니다…
i18n
redis
RedisInsight UI(i18next)의 국제화 규칙입니다. redisinsight/ui/** 아래의 사용자 표시 문자열을 추가하거나 변경할 때, ...를 편집할 때 사용합니다.
dead-dependencies
redis
RedisInsight에서 grep + leaf-check + build-gate 레시피를 사용하여 사용되지 않는("dead") npm 종속성을 찾아 안전하게 제거합니다. 종속성을 정리할 때 사용하세요,…