redis-semantic-cache

द्वारा redis

Redis LangCache का उपयोग Redis Cloud पर LLM प्रतिक्रियाओं के सिमैंटिक कैशिंग के लिए — SDK या REST API के माध्यम से search/set को कॉल करना, समानता सीमा को ट्यून करना,…

npx skills add https://github.com/redis/agent-skills --skill redis-semantic-cache

Redis Semantic Cache

Semantic caching for LLM responses with Redis Cloud's LangCache service. Stores prompts as embeddings; subsequent semantically-similar prompts return the cached response without re-calling the model.

LangCache is currently in preview on Redis Cloud. Features and behavior may change.

When to apply

  • Wrapping an LLM call (OpenAI, Anthropic, etc.) with a cache layer to cut cost and latency.
  • Caching RAG answers, classification outputs, or any deterministic LLM workload.
  • Tuning the precision/hit-rate trade-off for a semantic cache.
  • Splitting one application's LLM workloads across multiple cache instances.

1. The cache-aside flow

LangCache fits in front of any LLM call as a standard cache-aside pattern:

  1. Send the user's prompt to LangCache's search.
  2. Cache hit — return the stored response directly.
  3. Cache miss — call the LLM, then set the response so future similar prompts hit.
from langcache import LangCache
import os

lang_cache = LangCache(
    server_url=f"https://{os.getenv('HOST')}",
    cache_id=os.getenv("CACHE_ID"),
    api_key=os.getenv("API_KEY"),
)

result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.9)
if result:
    response = result[0]["response"]
else:
    response = llm.generate("What is Redis?")
    lang_cache.set(prompt="What is Redis?", response=response)

The same operations are available via REST (POST /v1/caches/{cacheId}/entries/search and POST /v1/caches/{cacheId}/entries) when an SDK isn't an option.

See references/langcache-usage.md for full SDK + REST samples and attribute-based storage.

2. Tune the similarity threshold

The threshold controls how close (in embedding cosine distance) a new prompt must be to a cached one to count as a hit. Higher = stricter match, fewer false positives. Lower = more hits, more risk of returning an off-topic answer.

ThresholdBehaviorUse when
0.95+Near-exact match requiredCustomer-facing answers where wrong responses are costly
0.9Balanced defaultMost workloads — start here
0.8Loose semantic matchInternal tools, exploratory queries, FAQ deduplication
# Stricter — fewer false positives
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.95)

# Looser — higher hit rate
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.8)

Adjust by watching the actual cache-hit rate and spot-checking that returned answers are still relevant.

See references/best-practices.md.

3. Separate caches per task type

Different LLM workloads should not share one cache — a "code question" prompt is semantically close to other code questions but has nothing to do with a password-reset support query, and crossing them returns garbage.

support_cache = LangCache(server_url=..., cache_id="support-cache-id", api_key=...)
code_cache    = LangCache(server_url=..., cache_id="code-cache-id",    api_key=...)

Create distinct cache IDs in Redis Cloud per task, and route each call to the right one. As a finer-grained alternative, store and search with custom attributes (e.g. {"category": "database"}) to keep tasks in the same cache but isolated by attribute filter — useful when the same prompt format spans subtopics.

References

redis की और Skills

docs-sync
redis
मास्टर ब्रांच के कार्यान्वयन और कॉन्फ़िगरेशन का विश्लेषण करें ताकि docs/, README.md, और प्रति-पैकेज README में गायब, गलत, या पुराने दस्तावेज़ का पता लगाया जा सके। उपयोग करें…
official
implement-command
redis
Add a new Redis command (or command variant) to node-redis end-to-end — the `<NAME>.ts` Command file, its registration with JSDoc in the package…
official
maintainer-review
redis
GitHub issue या pull request URL की समीक्षा node-redis अनुरक्षक के रूप में करें, जिसमें यह चरणबद्ध मूल्यांकन हो कि दावा वास्तविक है, व्यावहारिक रूप से महत्वपूर्ण है, पहले से...
official
pr-draft-summary
redis
node-redis के लिए आवश्यक PR-तैयार सारांश ब्लॉक, शाखा सुझाव, शीर्षक और ड्राफ्ट विवरण बनाएं। इसका उपयोग अंतिम प्रतिक्रिया से पहले किया जाना चाहिए जब भी...
official
runtime-behavior-probe
redis
रनटाइम व्यवहार जांच की योजना बनाएं और अस्थायी TypeScript प्रोब स्क्रिप्ट, सत्यापन मैट्रिक्स, स्थिति नियंत्रण और निष्कर्ष-प्रथम रिपोर्ट के साथ उन्हें क्रियान्वित करें। उपयोग करें…
official
backend
redis
NestJS बैकएंड विकास पैटर्न RedisInsight API के लिए: मॉड्यूल संरचना, सेवाएँ, नियंत्रक, DTO, निर्भरता इंजेक्शन, और त्रुटि प्रबंधन। उपयोग करें जब...
official
branches
redis
लोअरकेस केबब-केस का उपयोग करें, जिसमें टाइप प्रीफिक्स और इश्यू/टिकट आइडेंटिफायर हो। ब्रांच नाम GitHub Actions वर्कफ़्लो नियमों से मेल खाने चाहिए (.github/workflows/enforce-branch-name-rules.yml देखें)।
official
code-quality
redis
Code-quality standards for RedisInsight: TypeScript strictness, naming conventions (camelCase, PascalCase, UPPER_SNAKE_CASE), linting rules, no `any` without…
official