redis-semantic-cache

par redis

Guide Redis LangCache pour la mise en cache sémantique des réponses LLM sur Redis Cloud — appel de search/set via le SDK ou l'API REST, réglage du seuil de similarité,…

npx skills add https://github.com/redis/agent-skills --skill redis-semantic-cache

Redis Semantic Cache

Semantic caching for LLM responses with Redis Cloud's LangCache service. Stores prompts as embeddings; subsequent semantically-similar prompts return the cached response without re-calling the model.

LangCache is currently in preview on Redis Cloud. Features and behavior may change.

When to apply

  • Wrapping an LLM call (OpenAI, Anthropic, etc.) with a cache layer to cut cost and latency.
  • Caching RAG answers, classification outputs, or any deterministic LLM workload.
  • Tuning the precision/hit-rate trade-off for a semantic cache.
  • Splitting one application's LLM workloads across multiple cache instances.

1. The cache-aside flow

LangCache fits in front of any LLM call as a standard cache-aside pattern:

  1. Send the user's prompt to LangCache's search.
  2. Cache hit — return the stored response directly.
  3. Cache miss — call the LLM, then set the response so future similar prompts hit.
from langcache import LangCache
import os

lang_cache = LangCache(
    server_url=f"https://{os.getenv('HOST')}",
    cache_id=os.getenv("CACHE_ID"),
    api_key=os.getenv("API_KEY"),
)

result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.9)
if result:
    response = result[0]["response"]
else:
    response = llm.generate("What is Redis?")
    lang_cache.set(prompt="What is Redis?", response=response)

The same operations are available via REST (POST /v1/caches/{cacheId}/entries/search and POST /v1/caches/{cacheId}/entries) when an SDK isn't an option.

See references/langcache-usage.md for full SDK + REST samples and attribute-based storage.

2. Tune the similarity threshold

The threshold controls how close (in embedding cosine distance) a new prompt must be to a cached one to count as a hit. Higher = stricter match, fewer false positives. Lower = more hits, more risk of returning an off-topic answer.

ThresholdBehaviorUse when
0.95+Near-exact match requiredCustomer-facing answers where wrong responses are costly
0.9Balanced defaultMost workloads — start here
0.8Loose semantic matchInternal tools, exploratory queries, FAQ deduplication
# Stricter — fewer false positives
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.95)

# Looser — higher hit rate
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.8)

Adjust by watching the actual cache-hit rate and spot-checking that returned answers are still relevant.

See references/best-practices.md.

3. Separate caches per task type

Different LLM workloads should not share one cache — a "code question" prompt is semantically close to other code questions but has nothing to do with a password-reset support query, and crossing them returns garbage.

support_cache = LangCache(server_url=..., cache_id="support-cache-id", api_key=...)
code_cache    = LangCache(server_url=..., cache_id="code-cache-id",    api_key=...)

Create distinct cache IDs in Redis Cloud per task, and route each call to the right one. As a finer-grained alternative, store and search with custom attributes (e.g. {"category": "database"}) to keep tasks in the same cache but isolated by attribute filter — useful when the same prompt format spans subtopics.

References

Plus de skills de redis

docs-sync
redis
Analyser l'implémentation et la configuration de la branche master pour trouver la documentation manquante, incorrecte ou obsolète dans docs/, README.md et les READMEs par package. Utiliser…
redis-query-engine
redis
Conseils sur le moteur de requêtes Redis (RQE) couvrant la conception de schémas FT.CREATE, la sélection des types de champs (TEXT, TAG, NUMERIC, GEO, GEOSHAPE, VECTOR), la syntaxe des requêtes DIALECT 2,…
redis-search
redis
Conseils sur Redis Search couvrant la conception de schémas FT.CREATE, la sélection des types de champs (TEXT, TAG, NUMERIC, GEO, GEOSHAPE, VECTOR, chemin JSON), la syntaxe des requêtes DIALECT 2,…
redis-security
redis
Conseils de sécurité Redis couvrant l'authentification (requirepass et utilisateurs ACL), TLS, contrôle d'accès basé sur les ACL avec privilèges minimaux, restriction de l'exposition réseau via…
redis-vector-search
redis
Conseils sur la recherche vectorielle Redis couvrant le choix entre les algorithmes HNSW et FLAT, la configuration de l'index vectoriel (dimensions, métrique de distance, type de données), la recherche hybride filtrée…
bump-test-image
redis
Augmenter la version de l'image de test Docker Redis par défaut (redislabs/client-libs-test) dans le DEFAULT_DOCKER_CONFIG partagé et la matrice CI, puis force-push le…
i18n
redis
Conventions d'internationalisation pour l'interface RedisInsight (i18next). À utiliser lors de l'ajout ou de la modification de chaînes visibles par l'utilisateur sous redisinsight/ui/**, de la modification du…
dead-dependencies
redis
Trouver et supprimer en toute sécurité les dépendances npm inutilisées (« mortes ») dans RedisInsight à l'aide d'une recette grep + vérification des feuilles + gate de build. À utiliser lors du nettoyage des dépendances,…