redis-semantic-cache

द्वारा redis

Redis LangCache का उपयोग Redis Cloud पर LLM प्रतिक्रियाओं के सिमैंटिक कैशिंग के लिए — SDK या REST API के माध्यम से search/set को कॉल करना, समानता सीमा को ट्यून करना,…

npx skills add https://github.com/redis/agent-skills --skill redis-semantic-cache

Redis Semantic Cache

Semantic caching for LLM responses with Redis Cloud's LangCache service. Stores prompts as embeddings; subsequent semantically-similar prompts return the cached response without re-calling the model.

LangCache is currently in preview on Redis Cloud. Features and behavior may change.

When to apply

  • Wrapping an LLM call (OpenAI, Anthropic, etc.) with a cache layer to cut cost and latency.
  • Caching RAG answers, classification outputs, or any deterministic LLM workload.
  • Tuning the precision/hit-rate trade-off for a semantic cache.
  • Splitting one application's LLM workloads across multiple cache instances.

1. The cache-aside flow

LangCache fits in front of any LLM call as a standard cache-aside pattern:

  1. Send the user's prompt to LangCache's search.
  2. Cache hit — return the stored response directly.
  3. Cache miss — call the LLM, then set the response so future similar prompts hit.
from langcache import LangCache
import os

lang_cache = LangCache(
    server_url=f"https://{os.getenv('HOST')}",
    cache_id=os.getenv("CACHE_ID"),
    api_key=os.getenv("API_KEY"),
)

result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.9)
if result:
    response = result[0]["response"]
else:
    response = llm.generate("What is Redis?")
    lang_cache.set(prompt="What is Redis?", response=response)

The same operations are available via REST (POST /v1/caches/{cacheId}/entries/search and POST /v1/caches/{cacheId}/entries) when an SDK isn't an option.

See references/langcache-usage.md for full SDK + REST samples and attribute-based storage.

2. Tune the similarity threshold

The threshold controls how close (in embedding cosine distance) a new prompt must be to a cached one to count as a hit. Higher = stricter match, fewer false positives. Lower = more hits, more risk of returning an off-topic answer.

ThresholdBehaviorUse when
0.95+Near-exact match requiredCustomer-facing answers where wrong responses are costly
0.9Balanced defaultMost workloads — start here
0.8Loose semantic matchInternal tools, exploratory queries, FAQ deduplication
# Stricter — fewer false positives
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.95)

# Looser — higher hit rate
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.8)

Adjust by watching the actual cache-hit rate and spot-checking that returned answers are still relevant.

See references/best-practices.md.

3. Separate caches per task type

Different LLM workloads should not share one cache — a "code question" prompt is semantically close to other code questions but has nothing to do with a password-reset support query, and crossing them returns garbage.

support_cache = LangCache(server_url=..., cache_id="support-cache-id", api_key=...)
code_cache    = LangCache(server_url=..., cache_id="code-cache-id",    api_key=...)

Create distinct cache IDs in Redis Cloud per task, and route each call to the right one. As a finer-grained alternative, store and search with custom attributes (e.g. {"category": "database"}) to keep tasks in the same cache but isolated by attribute filter — useful when the same prompt format spans subtopics.

References

redis की और Skills

docs-sync
redis
मास्टर ब्रांच के कार्यान्वयन और कॉन्फ़िगरेशन का विश्लेषण करें ताकि docs/, README.md, और प्रति-पैकेज README में गायब, गलत, या पुराने दस्तावेज़ का पता लगाया जा सके। उपयोग करें…
redis-query-engine
redis
Redis Query Engine (RQE) मार्गदर्शन जिसमें FT.CREATE स्कीमा डिज़ाइन, फ़ील्ड प्रकार चयन (TEXT, TAG, NUMERIC, GEO, GEOSHAPE, VECTOR), DIALECT 2 क्वेरी सिंटैक्स शामिल है,…
redis-search
redis
Redis Search मार्गदर्शन जिसमें FT.CREATE स्कीमा डिज़ाइन, फ़ील्ड प्रकार चयन (TEXT, TAG, NUMERIC, GEO, GEOSHAPE, VECTOR, JSON path), DIALECT 2 क्वेरी सिंटैक्स शामिल है,…
redis-security
redis
Redis सुरक्षा मार्गदर्शन जिसमें प्रमाणीकरण (requirepass और ACL उपयोगकर्ता), TLS, ACL-आधारित न्यूनतम-विशेषाधिकार पहुँच नियंत्रण, नेटवर्क एक्सपोज़र को प्रतिबंधित करना शामिल है…
redis-vector-search
redis
Redis वेक्टर खोज मार्गदर्शन जिसमें HNSW बनाम FLAT एल्गोरिदम चयन, वेक्टर इंडेक्स कॉन्फ़िगरेशन (dims, दूरी मीट्रिक, डेटाटाइप), फ़िल्टर्ड हाइब्रिड खोज…
bump-test-image
redis
डिफ़ॉल्ट Redis डॉकर टेस्ट इमेज (redislabs/client-libs-test) को साझा DEFAULT_DOCKER_CONFIG और CI मैट्रिक्स में बढ़ाएँ, फिर फ़ोर्स-पुश करें…
i18n
redis
RedisInsight UI के लिए अंतर्राष्ट्रीयकरण परंपराएँ (i18next)। redisinsight/ui/** के अंतर्गत उपयोगकर्ता-सामने वाले स्ट्रिंग जोड़ते या बदलते समय उपयोग करें, संपादन करते समय…
dead-dependencies
redis
RedisInsight में grep + leaf-check + build-gate विधि का उपयोग करके अप्रयुक्त ("dead") npm निर्भरताओं को खोजें और सुरक्षित रूप से हटाएँ। निर्भरताओं की सफाई करते समय उपयोग करें,…