redis-semantic-cache

bởi redis

Hướng dẫn Redis LangCache về bộ nhớ đệm ngữ nghĩa cho phản hồi LLM trên Redis Cloud — gọi search/set qua SDK hoặc REST API, điều chỉnh ngưỡng tương đồng,…

npx skills add https://github.com/redis/agent-skills --skill redis-semantic-cache

Redis Semantic Cache

Semantic caching for LLM responses with Redis Cloud's LangCache service. Stores prompts as embeddings; subsequent semantically-similar prompts return the cached response without re-calling the model.

LangCache is currently in preview on Redis Cloud. Features and behavior may change.

When to apply

  • Wrapping an LLM call (OpenAI, Anthropic, etc.) with a cache layer to cut cost and latency.
  • Caching RAG answers, classification outputs, or any deterministic LLM workload.
  • Tuning the precision/hit-rate trade-off for a semantic cache.
  • Splitting one application's LLM workloads across multiple cache instances.

1. The cache-aside flow

LangCache fits in front of any LLM call as a standard cache-aside pattern:

  1. Send the user's prompt to LangCache's search.
  2. Cache hit — return the stored response directly.
  3. Cache miss — call the LLM, then set the response so future similar prompts hit.
from langcache import LangCache
import os

lang_cache = LangCache(
    server_url=f"https://{os.getenv('HOST')}",
    cache_id=os.getenv("CACHE_ID"),
    api_key=os.getenv("API_KEY"),
)

result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.9)
if result:
    response = result[0]["response"]
else:
    response = llm.generate("What is Redis?")
    lang_cache.set(prompt="What is Redis?", response=response)

The same operations are available via REST (POST /v1/caches/{cacheId}/entries/search and POST /v1/caches/{cacheId}/entries) when an SDK isn't an option.

See references/langcache-usage.md for full SDK + REST samples and attribute-based storage.

2. Tune the similarity threshold

The threshold controls how close (in embedding cosine distance) a new prompt must be to a cached one to count as a hit. Higher = stricter match, fewer false positives. Lower = more hits, more risk of returning an off-topic answer.

ThresholdBehaviorUse when
0.95+Near-exact match requiredCustomer-facing answers where wrong responses are costly
0.9Balanced defaultMost workloads — start here
0.8Loose semantic matchInternal tools, exploratory queries, FAQ deduplication
# Stricter — fewer false positives
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.95)

# Looser — higher hit rate
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.8)

Adjust by watching the actual cache-hit rate and spot-checking that returned answers are still relevant.

See references/best-practices.md.

3. Separate caches per task type

Different LLM workloads should not share one cache — a "code question" prompt is semantically close to other code questions but has nothing to do with a password-reset support query, and crossing them returns garbage.

support_cache = LangCache(server_url=..., cache_id="support-cache-id", api_key=...)
code_cache    = LangCache(server_url=..., cache_id="code-cache-id",    api_key=...)

Create distinct cache IDs in Redis Cloud per task, and route each call to the right one. As a finer-grained alternative, store and search with custom attributes (e.g. {"category": "database"}) to keep tasks in the same cache but isolated by attribute filter — useful when the same prompt format spans subtopics.

References

Thêm skills từ redis

docs-sync
redis
Phân tích triển khai và cấu hình của nhánh master để tìm tài liệu bị thiếu, sai hoặc lỗi thời trong docs/, README.md và các README theo từng gói. Sử dụng…
official
implement-command
redis
Add a new Redis command (or command variant) to node-redis end-to-end — the `<NAME>.ts` Command file, its registration with JSDoc in the package…
official
maintainer-review
redis
Xem xét một URL issue hoặc pull request trên GitHub với tư cách là người bảo trì node-redis, đánh giá theo từng giai đoạn xem tuyên bố đó có thực tế, quan trọng trong thực tế, đã được…
official
pr-draft-summary
redis
Tạo khối tóm tắt sẵn sàng cho PR, gợi ý nhánh, tiêu đề và mô tả dự thảo cho node-redis. Phải được sử dụng trước phản hồi cuối cùng bất cứ khi nào…
official
runtime-behavior-probe
redis
Lập kế hoạch và thực hiện các cuộc điều tra hành vi thời gian chạy với các tập lệnh thăm dò TypeScript tạm thời, ma trận xác thực, kiểm soát trạng thái và báo cáo ưu tiên phát hiện. Sử dụng…
official
backend
redis
Các mẫu phát triển backend NestJS cho API RedisInsight: cấu trúc module, dịch vụ, bộ điều khiển, DTO, tiêm phụ thuộc và xử lý lỗi. Sử dụng khi…
official
branches
redis
Sử dụng chữ thường, kiểu kebab-case với tiền tố loại và mã định danh issue/ticket. Tên nhánh phải khớp với quy tắc của quy trình GitHub Actions (xem .github/workflows/enforce-branch-name-rules.yml).
official
code-quality
redis
Code-quality standards for RedisInsight: TypeScript strictness, naming conventions (camelCase, PascalCase, UPPER_SNAKE_CASE), linting rules, no `any` without…
official