Revoked Access and Cached AI Answers

A previously authorized RAG answer is not permanently safe to reuse. Review permission changes across retrieval, answer caches, history and in-flight requests.

A cache hit still needs an authorization decision

When access to a source is revoked, an AI application must not keep serving its restricted information just because an answer was generated earlier. Treat cached answers, conversation history and assembled context as derived content with access requirements. A past authorization decision does not establish permission for a new request.

The difficult case is not only a stale search index. An answer cache can bypass retrieval completely. A conversation can contain a summary of a removed document. An in-flight response can finish after a permission change. Protecting the next search query leaves those paths unresolved.

This article proposes a review method for retrieval-augmented generation systems. Its project-transfer example is hypothetical, not a customer security incident. The method concerns future server-controlled disclosure. It cannot erase information that a previously authorized reader already saw, copied or downloaded, and it should not promise that capability.

Start with the question your security owner can actually answer: after a revocation is effective, which paths may still deliver protected content, under what consistency assumptions, and for how long? Make the answer explicit enough to test. “Permissions are supported” is not a usable revocation contract.

Inventory the copies that retrieval creates

List the source document, extracted text, indexed chunks, retrieval results, assembled prompt context, generated answers, conversation records and exports. Add background jobs and operational traces where they contain content. Each copy has a different lifecycle and may have a different serving endpoint.

A source deletion or ACL update does not necessarily update every derived artifact at once. Record which event reaches each store, how missed events are detected and how serving is blocked while an update is unresolved. The inventory should name an owner and a verification method, not just a database technology.

Separate access revocation from deletion. A document can remain valid for another team while one reader loses access. Removing the whole document may unnecessarily affect other users. Conversely, hiding a search result from one user does not satisfy a separate requirement to remove derived data from storage.

Include the history endpoint in this review. If it returns old messages containing restricted excerpts, the system is still serving protected content even when the new-answer endpoint is correct. Decide whether history requires source-level reauthorization, restricted-message replacement or a different documented retention and viewing policy.

Enforce access with trusted identity, not a submitted filter

The OWASP Authorization Cheat Sheet recommends denying by default, validating permission on every request and enforcing checks outside the client. For a RAG application, apply those principles to cached answers and history as well as fresh retrieval. A cache key is not an authorization policy.

Construct tenant and principal scope from trusted application identity. Do not accept a browser-supplied tenant ID or group list as proof of membership. The application needs to know which identity assertions are current and how permission changes invalidate or supersede older assertions.

Microsoft's Azure AI Search security-filter pattern illustrates filtering indexed documents by principal identifiers. It also explicitly distinguishes that string-filter mechanism from authentication and authorization. The surrounding application remains responsible for supplying trustworthy scope and applying the control consistently.

Test direct object access, not only the chat route. A guessed answer identifier, citation URL, export link or conversation identifier must not let a reader bypass the same boundary. Obscure identifiers can reduce accidental discovery; they do not establish entitlement to the object behind them.

Worked example: a project transfer during answer generation

Consider a hypothetical engineer who can read Project Orion documents. At 10:00, the engineer asks for a summary and generation starts. At 10:01, the project owner removes the engineer's access. At 10:02, the generated answer is ready. The example exposes a race between eligibility at retrieval and eligibility at delivery.

The application needs a defined decision point. If its contract requires current authorization when delivering a new answer, it must recheck the relevant access state before release. Where the store supports it, a revision or authorization epoch can help detect that the earlier decision became stale. The implementation must still address changes between the final check and delivery.

| Serving path | Why an earlier check is insufficient | Acceptance test | | --- | --- | --- | | New retrieval request | Index ACL metadata may lag | Revoked reader cannot obtain affected chunks | | Existing cached answer | Retrieval may not execute | Cache hit passes current entitlement checks | | In-flight generation | Access changes after retrieval | Changed entitlement blocks release under the chosen contract | | Conversation history | Old message contains restricted information | History route enforces the defined viewing policy | | Citation or export | Separate endpoint may omit checks | Direct request cannot bypass object authorization |

If the authorization service is unavailable, do not turn an unknown entitlement into permission by serving the old cache entry. Use a defined unavailable or access-check-pending state for protected content. Public material may have a separate path, but the application must establish that classification rather than guess it from a familiar query.

The exact consistency mechanism depends on the system. An epoch check alone is not an atomic solution across unrelated services. Document the enforcement point, propagation behavior and residual race window, then verify them with the security owner. Avoid an “instant revocation everywhere” claim unless the implementation can demonstrate it.

Bind answer reuse to evidence and current scope

Retain provenance sufficient to identify which protected sources contributed to an answer. A response made from several sources can inherit restrictions from more than one document. If the application cannot reliably determine eligibility for that derived answer, refusing reuse is safer than assuming the most permissive source controls it.

A practical cache entry can record tenant scope, source revisions, applicable access-policy revision, creation time and task configuration. Those fields support invalidation and evaluation. They are not secrets that replace authorization, and they do not authorize a caller merely because the caller knows their values.

Recheck access before replaying cached content. If the entitlement snapshot is stale, resolve current scope or reject the protected response according to policy. Choose the permissible freshness window deliberately. A short time-to-live limits age; it does not prove that every answer remains permitted until expiry.

Semantic matching introduces a second risk. A similar question from another user must not retrieve an answer prepared under broader permissions. Keep similarity ranking inside the already established eligibility boundary. Do not search a shared response pool first and then hope a text-level redaction will remove every restricted fact.

Close invalidation races and background-work gaps

Revocation events can arrive late, twice or out of order. Use a durable revision model where available, reject stale updates and make recovery observable. Track the permission revision applied by each serving store. A successful event delivery is weaker evidence than a readback showing the intended restriction is active.

Also test an in-flight cache fill. An invalidation can remove an entry while an older request is still generating it. Without a guard, that request may recreate the restricted entry immediately afterwards. Compare the relevant revision before accepting the fill, and keep serving authorization independent of whether invalidation completed.

Background summaries and scheduled exports need current scope too. A job created before revocation should not automatically run with the requester's old access. Preserve the job identity, re-evaluate authority at the defined execution boundary and prevent a stale job from creating a new unrestricted copy.

Avoid storing sensitive answer bodies in broad observability systems just to debug propagation. Record object identifiers, decision revisions, timing and allow or deny outcomes where sufficient. If deeper content inspection is necessary, use an explicitly protected investigation path. Logging should help verify the boundary, not create an untracked alternate source of the content.

Limitations: revocation does not undo prior disclosure

Server-side controls can restrict future access to resources the system controls. They cannot reliably retract a screenshot, a copied paragraph or an already downloaded file. Distinguish revocation latency from historical disclosure and retention requirements when explaining the design to stakeholders.

Streaming requires a separate decision. Once a protected token has reached the reader, a later permission check cannot recall it. Choose how streams respond to mid-stream changes, what consistency guarantee is feasible and whether sensitive tasks should buffer output before release. That choice has latency and complexity costs, not just a security checkbox.

Removing a cited paragraph from an answer is not automatically sufficient. The rest of the summary may reveal its content indirectly. Where provenance is incomplete, rebuild from currently permitted sources or withhold the answer. Do not describe automated redaction as complete protection without evidence for the actual task and data.

Build a revocation test that exercises every serving path

Create a fixture with two users in the same tenant, different project access and one restricted source. Generate an answer while access is valid. Revoke one user's access, then request the cached answer, the conversation, the citation and the export directly. Assert that the other user's permitted access continues working where the policy allows it.

Repeat with delayed ACL propagation, a missing invalidation event, an in-flight fill and an unavailable authorization dependency. Capture when the change became authoritative and when each serving path enforced it. Use those timestamps to evaluate the documented contract, rather than measuring only whether the cache entry disappeared.

Read stale policy answers in a semantic cache for content-version changes and retrieval changes that break answers for regression fixtures. The AI change-release playbook helps organize the evidence. For implementation help, explore production AI systems or share an access-control question. Reading these resources does not require sharing personal details.