60-sec Short LLM Apps: Caching & RAG 45 sec Failure mode + fix

Post-Filter RAG Leaks

Your RAG bot just leaked salaries.

Video premiering soon on @AI.JoinDev Subscribe to get it first

Why it leaks

Filtering after retrieval means the LLM already read the forbidden text. Any answer can quote it.

The fix

Pre-filter: apply the user's ACL inside the vector search, before anything reaches the LLM.

Context

A company-wide RAG assistant answers employee questions over documents from every team (HR, finance, engineering). All teams' chunks live in one shared vector index. Each document has an access-control list (ACL); each employee should only ever get answers grounded in documents they are allowed to open.

Architecture

ComponentRoleNotes
Employee (client)Sends a natural-language questionAuthenticated; identity + groups known
Shared RetrieverEmbeds the query, runs top-k vector searchNo per-user filter in the search call
Vector IndexStores chunks + embeddings for all teamsChunks carry ACL metadata, but it isn't used at query time
LLMGenerates the answer from the retrieved contextSees every chunk it is given
Permission CheckFilters or redacts after generation"Post-filter" — runs on the output / cited sources

Request flow

  1. Employee asks: "Show me Q3 salaries".
  2. Retriever searches the shared index and returns top-k chunks by similarity only.
  3. Top-k contains an allowed chunk (public comp policy) and a forbidden one (HR salary table).
  4. Both chunks are placed in the LLM prompt as context.
  5. The LLM writes an answer grounded in everything it saw.
  6. A permission check then drops citations the user can't open — but the text of the answer stays.

The failure mode

The permission check runs on the output, but the leak happened at the input. Once a forbidden chunk is in the context window, the model can quote it, paraphrase it, or aggregate it ("the average engineer salary is …"). Removing the citation or the source link does not remove the information from the generated text. It is easy to miss because the demo looks correct: forbidden documents never appear as sources in the UI.

Why it happens

Authorization is applied at the wrong layer: after retrieval and generation instead of as a constraint on retrieval.

The fix

Pre-filter: pass the user's identity/groups into the vector search and filter on ACL metadata inside the query (metadata filter, per-tenant namespace/partition, or row-level security in the store). Forbidden chunks are never retrieved, so they can never reach the prompt. Keep the post-check as defense in depth, not as the control.

Trade-offs

  • Filtered ANN search can lower recall when the filter is very selective — raise k/ef_search, or use per-tenant partitions for large tenants.
  • ACLs must be synced into the index (and re-synced on permission changes); stale ACLs are a new failure mode.
  • Per-tenant namespaces isolate cleanly but cost more memory and operational overhead.

Numbers worth knowing

  • Illustrative: top-k = 8 over a shared index means one question can pull chunks from several teams at once.

Coming next in the series: Pre-Filter RAG

Found this useful?

Subscribe for the next episode, or share it with the person who owns this part of your stack.

Keep going