Deep dive LLM Apps: Caching & RAG 6 min Failure mode + fix

Semantic Caching for LLMs

Your AI cache just answered the wrong customer.

How it works, why it lies, and how it leaks

Video premiering soon on @AI.JoinDev Subscribe to get it first

Chapters

  1. Intro
  2. Why cache LLM calls
  3. How a semantic cache works
  4. The false-hit trap
  5. The cross-tenant leak
  6. The fix
  7. Recap

Who it's for

Backend and platform engineers putting an LLM gateway in front of a model to cut latency and cost, especially in multi-tenant SaaS assistants.

Context

A customer-support or internal-knowledge assistant behind an LLM gateway. Many users ask the same things in different words, so the team adds a semantic cache: embed each question, look up the nearest earlier question in a vector index, and reuse its answer when the similarity clears a threshold. In a multi-tenant product, one deployment serves many companies.

Architecture

ComponentRoleNotes
LLM gatewaySingle entry point for model callsOwns caching, routing, rate limits
EmbedderTurns the question into a vectorSame model for writes and reads
Vector cacheNearest-neighbour search over past questionsCosine similarity, threshold (e.g. 0.90)
LLMGenerates an answer on a missThe expensive, slow path
Hit verifier (fix)Confirms a candidate hit has the same intentCross-encoder re-ranker or a small model
Scoped key / partition (fix)Limits search to what this caller may seeTenant + role hash + model + prompt version

Request flow

  1. The user's question reaches the gateway.
  2. The gateway embeds it and queries the vector cache for the nearest stored question.
  3. If the score clears the threshold, the stored answer is returned (a hit, tens of milliseconds).
  4. Otherwise the model answers (a miss, seconds), and the question vector and answer are stored.

The failure modes

False hit. "How do I cancel my plan?" and "How do I cancel my order?" differ by one word, so their embeddings sit very close (similarity ~0.94, illustrative). The cache returns the order-refund answer to a subscription question. Nothing errors, nothing is logged, and the answer is delivered confidently and fast.

Cross-tenant leak. Company A's user asks for Q3 revenue; on a miss the model answers from company A's private data and the answer is cached. Company B's user asks almost the same question, the embeddings match, and company B receives company A's numbers.

Why it happens

The cache key is only the embedding of the question. It captures what was asked, not who asked, what they may see, or which model and prompt produced the answer, and similarity is treated as sameness.

The fix

  • Scope the key: tenant id, a hash of the caller's roles/ACL, model version and prompt version. Search only that partition (or one index per tenant).
  • Measure the threshold: label a few hundred real query pairs as same/different intent and pick the threshold where the false-hit rate is acceptable for the domain.
  • Verify hits: treat a hit as a candidate and confirm intent with a cross-encoder re-ranker or a small model before returning it.
  • Don't share personal answers: answers built from one user's data are cached per user, or not at all.
  • Expire and purge: TTLs, plus purges when the source documents change.
  • Observe: log hit scores and sample hits for human review so false hits show up on a dashboard, not in a complaint.

Trade-offs

  • Scoping the key lowers the hit rate: the same public question is cached once per tenant/role set.
  • A verifier adds latency (a re-ranker call) and cost to every hit, though far less than a full model call.
  • A stricter threshold means fewer false hits but fewer hits overall; the right point depends on the cost of a wrong answer.
  • Per-tenant indexes add operational overhead (many small indexes, eviction per tenant).

Numbers worth knowing

  • All latency, hit-rate and false-hit figures in the video are illustrative (1.8 s model call, 40 ms hit, 30% repeat questions, the threshold curve).
  • Open-source semantic caches: GPTCache (Zilliz) and RedisVL's SemanticCache; both expose a similarity/distance threshold you must tune.

The fix in code

semantic_cache.py
def cache_key(req):
    return {
        "tenant": req.tenant_id,
        "acl": hash_roles(req.user.roles),
        "model": MODEL_VERSION,
        "prompt": PROMPT_VERSION,
    }

def lookup(req, vector):
    hit = index(cache_key(req)).nearest(vector)
    if not hit or hit.score < THRESHOLD:
        return None  # miss: call the model
    if not same_intent(req.text, hit.question):
        return None  # false hit caught
    return hit.answer

Production checklist

Scope every key

Tenant, roles, model and prompt version

Measure the threshold

Labelled query pairs from real traffic

Verify hits

Re-rank or ask a small model

Skip personal answers

Never share answers built from one user's data

Expire and purge

TTLs, plus purge when source docs change

Watch hit quality

Log scores and sample hits for review

Recap: before and after

Naive semantic cache

  • Key is the embedding only
  • Threshold picked by feel
  • Every hit trusted
  • One index for all tenants
vs

Production semantic cache

  • Key scoped to tenant and roles
  • Threshold measured on real pairs
  • Hits verified before use
  • Partitions, TTLs and purges

A semantic cache is only as safe as its key.

Sources

  • GPTCache project documentation (Zilliz) — semantic cache architecture and similarity evaluation.

  • Redis documentation, RedisVL SemanticCache — distance threshold and per-user/tenant filtering.

Coming next in the series: Prompt Injection Through Tools

Found this useful?

Subscribe for the next episode, or share it with the person who owns this part of your stack.

Keep going