Semantic Caching for LLMs
Your AI cache just answered the wrong customer.
How it works, why it lies, and how it leaks
Chapters
- Intro
- Why cache LLM calls
- How a semantic cache works
- The false-hit trap
- The cross-tenant leak
- The fix
- Recap
Who it's for
Backend and platform engineers putting an LLM gateway in front of a model to cut latency and cost, especially in multi-tenant SaaS assistants.
Context
A customer-support or internal-knowledge assistant behind an LLM gateway. Many users ask the same things in different words, so the team adds a semantic cache: embed each question, look up the nearest earlier question in a vector index, and reuse its answer when the similarity clears a threshold. In a multi-tenant product, one deployment serves many companies.
Architecture
| Component | Role | Notes |
|---|---|---|
| LLM gateway | Single entry point for model calls | Owns caching, routing, rate limits |
| Embedder | Turns the question into a vector | Same model for writes and reads |
| Vector cache | Nearest-neighbour search over past questions | Cosine similarity, threshold (e.g. 0.90) |
| LLM | Generates an answer on a miss | The expensive, slow path |
| Hit verifier (fix) | Confirms a candidate hit has the same intent | Cross-encoder re-ranker or a small model |
| Scoped key / partition (fix) | Limits search to what this caller may see | Tenant + role hash + model + prompt version |
Request flow
- The user's question reaches the gateway.
- The gateway embeds it and queries the vector cache for the nearest stored question.
- If the score clears the threshold, the stored answer is returned (a hit, tens of milliseconds).
- Otherwise the model answers (a miss, seconds), and the question vector and answer are stored.
The failure modes
False hit. "How do I cancel my plan?" and "How do I cancel my order?" differ by one word, so their embeddings sit very close (similarity ~0.94, illustrative). The cache returns the order-refund answer to a subscription question. Nothing errors, nothing is logged, and the answer is delivered confidently and fast.
Cross-tenant leak. Company A's user asks for Q3 revenue; on a miss the model answers from company A's private data and the answer is cached. Company B's user asks almost the same question, the embeddings match, and company B receives company A's numbers.
Why it happens
The cache key is only the embedding of the question. It captures what was asked, not who asked, what they may see, or which model and prompt produced the answer, and similarity is treated as sameness.
The fix
- Scope the key: tenant id, a hash of the caller's roles/ACL, model version and prompt version. Search only that partition (or one index per tenant).
- Measure the threshold: label a few hundred real query pairs as same/different intent and pick the threshold where the false-hit rate is acceptable for the domain.
- Verify hits: treat a hit as a candidate and confirm intent with a cross-encoder re-ranker or a small model before returning it.
- Don't share personal answers: answers built from one user's data are cached per user, or not at all.
- Expire and purge: TTLs, plus purges when the source documents change.
- Observe: log hit scores and sample hits for human review so false hits show up on a dashboard, not in a complaint.
Trade-offs
- Scoping the key lowers the hit rate: the same public question is cached once per tenant/role set.
- A verifier adds latency (a re-ranker call) and cost to every hit, though far less than a full model call.
- A stricter threshold means fewer false hits but fewer hits overall; the right point depends on the cost of a wrong answer.
- Per-tenant indexes add operational overhead (many small indexes, eviction per tenant).
Numbers worth knowing
- All latency, hit-rate and false-hit figures in the video are illustrative (1.8 s model call, 40 ms hit, 30% repeat questions, the threshold curve).
- Open-source semantic caches: GPTCache (Zilliz) and RedisVL's
SemanticCache; both expose a similarity/distance threshold you must tune.
The fix in code
def cache_key(req):
return {
"tenant": req.tenant_id,
"acl": hash_roles(req.user.roles),
"model": MODEL_VERSION,
"prompt": PROMPT_VERSION,
}
def lookup(req, vector):
hit = index(cache_key(req)).nearest(vector)
if not hit or hit.score < THRESHOLD:
return None # miss: call the model
if not same_intent(req.text, hit.question):
return None # false hit caught
return hit.answer
Production checklist
Scope every key
Tenant, roles, model and prompt version
Measure the threshold
Labelled query pairs from real traffic
Verify hits
Re-rank or ask a small model
Skip personal answers
Never share answers built from one user's data
Expire and purge
TTLs, plus purge when source docs change
Watch hit quality
Log scores and sample hits for review
Recap: before and after
Naive semantic cache
- Key is the embedding only
- Threshold picked by feel
- Every hit trusted
- One index for all tenants
Production semantic cache
- Key scoped to tenant and roles
- Threshold measured on real pairs
- Hits verified before use
- Partitions, TTLs and purges
A semantic cache is only as safe as its key.
Sources
GPTCache project documentation (Zilliz) — semantic cache architecture and similarity evaluation.
Redis documentation, RedisVL
SemanticCache— distance threshold and per-user/tenant filtering.
Coming next in the series: Prompt Injection Through Tools