Permission Changes in Enterprise AI Search
Choose a revocation contract for enterprise AI search. Compare synchronized permissions, current checks and release controls across caches and running answers.
audience="Application owners, retrieval engineers and security reviewers selecting how an enterprise assistant handles changed access." decision="Which policy observation permits disclosure, what happens to already assembled answers and how unresolved permission propagation affects service." position="Select an explicit disclosure contract for each source class. Treat retrieval eligibility, derived-answer reuse and response release as separate boundaries with evidence of current scope." scope="A proposed design framework using a hypothetical project-document assistant. Not a deployed customer architecture, security certification or promise of instantaneous revocation." outputs={['A source-class revocation contract', 'A serving-surface and policy-source inventory', 'An option comparison with consistency limits', 'A derived-answer provenance specification', 'An in-flight response disposition', 'An acceptance and change-review record']} />
Executive summary
An enterprise assistant can honor permissions when it retrieves a document and still disclose that document after access changes. The answer may already exist in a response cache. A conversation may contain a summary. A background request may have fetched the source before the change and release its answer afterward. None of these paths necessarily makes another search request. The architectural decision is therefore broader than choosing a search filter: which observation of permission is sufficient at each server-controlled disclosure boundary, and what must happen when that observation is stale or unavailable?
Thesis: define the revocation contract before selecting the consistency mechanism. A synchronized search index, a current authorization check and a response-release control establish different things. Use them together only where the required boundary and observable implementation justify the combination. Do not label a system “instant revocation” because its identity provider accepted a change or because its next clean-session search returned no result. Record the policy state used by each path and the conditions under which an answer may be released, reused, withheld or recomputed.
The recommended starting point is one protected source class, one ordinary role-removal path and a synthetic answer that has identifiable provenance. Compare the normal serving path with a warmed cache, existing history and a delayed response. Include a reader who retains access so the design cannot pass merely by denying everyone. Use the results to choose the operational boundary, not to manufacture a green security score. This paper develops the design choices; the revocation test playbook provides an executable exercise once the contract has been agreed.
Define disclosure, revocation and the effective boundary
For this framework, disclosure means the application newly serves protected information through a controlled interface. It includes answer text, excerpts, document previews, citation labels and downloadable reports where those reveal restricted facts. Revocation means a previously permitted subject-resource relationship is removed or superseded under the chosen policy. Effective boundary means the event after which the application must apply the new restriction to a specified serving path. These are design definitions, not a legal interpretation or a claim that every provider uses the same terms.
Separate the requested change, its authoritative acceptance and the serving system's observation. An administrator can submit a group change that fails. A policy store can accept it while a search connector has not applied it. An application can observe it while an old response is already on its way to a browser. Record these milestones independently. Where the business chooses a bounded synchronization interval, say which interval is measured and how the application behaves if the bound is exceeded. A tolerated test delay is not automatically an approved policy.
The contract must also state what it cannot do. A reader who legitimately downloaded a report yesterday may retain that copy after today's revocation. Preventing a new server download does not erase an external copy. Previously delivered tokens cannot be recalled by hiding the remainder of a stream. Those limits do not make revocation useless; they define its actual control surface. Avoid wording that implies removal from memory, screenshots or unmanaged devices unless a separate mechanism and authority genuinely support that scope.
Use concrete subject and object identities. A project label may cover several sources, while one document revision can have a different restriction from its replacement. Tenant membership, group membership, source ACLs and temporary exceptions can interact. The design record needs the relevant relationship, revision and effective event, not just an “employee removed” flag. Start with the relationship the source owner can explain and expand only after the application can preserve that meaning through its derived representations.
A project transfer reveals three independent decisions
Consider a hypothetical assistant serving project planning documents. Mira initially has access to Project North. She asks for a summary, the application retrieves two protected revisions and generation begins. A project owner then removes Mira's membership while another colleague retains access. The answer is ready after the removal. Separately, yesterday's answer is in a semantic cache and an earlier conversation contains a protected excerpt. This scenario is illustrative, not an Ampity customer incident or a description of a particular vendor deployment.
The first decision is retrieval eligibility: which sources could be selected when the request assembled context? The second is derived-content reuse: may the application return an answer produced under an earlier entitlement? The third is release eligibility: may a response already being generated be delivered after the effective change? Treating these as one check obscures the actual race. Retrieval can be correct at its own observation point while response release violates the agreed boundary. Cache reuse can bypass both fresh retrieval and generation.
State the desired treatment of Mira's request before discussing implementation. The owner might require withholding any not-yet-released answer once the new policy is effective. Another source class might allow a request admitted before a defined boundary to finish. These are materially different confidentiality commitments. Neither should be selected implicitly by an SDK's buffering behavior. A product team must explain the experience and a security owner must accept the boundary, including any residual race or stale-policy interval that the implementation cannot eliminate.
Preserve the authorized colleague as a control. If both readers lose the document because the index was deleted, the design has not demonstrated per-reader revocation. If the assistant continues serving old excerpts to Mira because it only checks new searches, the design has not demonstrated coverage of derived content. The useful artifact is a result for each subject and serving surface, with the observed policy and response disposition. A single “RAG access control enabled” checkbox cannot capture those differences.
Establish a trusted policy source and scope
Choose the authority for each permission relationship. It may be the source repository, an application policy store or an identity system combined with source ACLs. Document how conflicting observations are resolved. A browser-provided group list is not an authority merely because it has the right field name. A connector's copied ACL is evidence of what it last synchronized, not necessarily the source's present decision. Establish who can change the relationship, how revisions are identified and which trusted observation the serving layer may consume.
OWASP's authorization guidance recommends denying by default, checking permissions on every request and placing enforcement in the appropriate server-side boundary. Apply that principle to actual objects and endpoints, not only the visible chat interface. A direct answer identifier, report URL or source-preview route can disclose the same information through another handler. This recommendation does not prescribe one identity vendor or prove that a chosen implementation is correct; it makes omitted serving paths part of the review.
Scope the service identity separately from the reader identity. A backend may have broad technical permission to query an index, while the person asking a question has narrow document access. Passing the backend's permission through as if it represented the reader can defeat the intended boundary. Name both actors in the design and show where the reader's scope constrains retrieval and delivery. Administrative investigation paths need their own authority, evidence and retention controls; they should not become an ordinary assistant fallback.
Account for relationship expansion. A source can inherit access from a parent folder, a team can inherit membership from another group, and a temporary grant can expire. A simple comparison of one stored group string may be insufficient for that model. Either implement the needed relationships or limit the supported policy honestly. Record unsupported nesting and inheritance as source-onboarding constraints. Do not ingest a source whose permission semantics the application cannot represent, then compensate with a general instruction telling the model not to reveal it.
Option one: synchronized permission metadata
In a synchronized design, the serving index stores permission metadata alongside searchable content. Queries use trusted reader scope to exclude ineligible records. This can make retrieval efficient and preserve source-specific policy information near each chunk. Its central limitation is the gap between a source change and the index applying that change. The design must identify the synchronization mechanism, its coverage and the failure behavior when an update is missed. Scheduled ingestion alone does not establish a maximum revocation delay.
Azure AI Search's document-access overview distinguishes generally available security-string filters from native token-based permission approaches marked preview. The documented query enforcement uses permission metadata already synchronized into the index. As inspected in October 2026, the native patterns and their source-specific synchronization requirements should not be described as universally available production guarantees. Check the actual API, source connector and permissions supported by the chosen deployment rather than transferring a feature label from one scenario to another.
The decision record should say how stale metadata affects serving. One proposed approach is to block an affected source family until a required permission revision is applied and observed. Another is to allow an explicitly approved freshness interval, then stop protected retrieval if the observation falls outside it. These are application design choices, not undocumented vendor guarantees. Evaluate whether the system has enough revision information to make either choice meaningful. A timestamp saying a job ran is weaker than a readback tied to the permission change.
Synchronized metadata can be appropriate where the organization accepts its measured update boundary and the index's policy representation matches the source. It is not sufficient by itself for a requirement that forbids every new disclosure immediately after a source change, especially when old answers bypass retrieval. Document that mismatch before selecting the option. A more frequent job may reduce observed delay, but it does not automatically solve event loss, inherited changes or a serving cache that still trusts yesterday's eligibility.
Option two: current checks at serving boundaries
A current-check design consults an authorization authority when deciding whether a candidate source or derived answer may be served. It reduces reliance on old permission metadata when the authority can answer with the required freshness. It also creates an operational dependency: protected serving now needs a usable policy decision. The application must define timeout, unavailable and ambiguous states without turning them into permission. A cached allow response can improve latency, but doing so reintroduces a freshness contract that must be explicitly evaluated.
OpenFGA's consistency documentation describes MINIMIZE_LATENCY, which can use enabled caches, and HIGHER_CONSISTENCY, which skips those caches for a database query. The page also states that caching is disabled by default and discusses the performance consequences of higher-consistency requests. Do not describe the option as an atomic transaction covering an unrelated search index, generation job and response socket. Selecting a query mode changes the authorization query's behavior; application release coordination remains a separate responsibility.
Estimate the check workload using the actual serving pattern. One answer may cite several sources, and one page may request history, previews and a report. Decide which checks can be batched or safely shared under the chosen contract, and which must remain object-specific. Measure normal and dependency-failure latency rather than assuming an authority can absorb every expansion cheaply. The relevant budget is the cost and availability of accepted responses with correct scope, not the fastest unauthorized shortcut through an old cache.
Current checks are useful when policy observations are authoritative enough for the required disclosure decision and all relevant serving paths can use them. They are not sufficient when lineage is missing, when a history message combines unidentifiable sources or when the application cannot coordinate the check with release. State those prerequisites clearly. If an answer cannot be proven eligible, refuse its reuse or recompute from eligible material. Do not ask the model to infer permission from the tone or apparent sensitivity of its own output.
Option three: explicit release coordination
A release-control design identifies the point at which protected output crosses from held application state into an externally visible response. It can associate pending work with the subject, contributing sources and the relevant policy observation. When a change affects that work, the application can withhold, cancel or re-evaluate it according to the contract. The design is especially important for delayed jobs and reused answers because their original admission happened before the present request to disclose them.
The Zanzibar research paper describes authorization decisions that respect causal ordering of permission and object-content changes in Google's system. It is a useful example of why ordering belongs in the design rather than being reduced to cache age. It is not proof that another product, a generic epoch field or this proposed assistant inherits those guarantees. An implementation claiming a comparable boundary needs evidence of its own coordination protocol, supported operations and failure behavior.
Avoid the misleading claim that checking an epoch immediately before writing a response solves every race. If the check and release happen in different services, a change can occur between them. If the application streams bytes after the check, an update can become effective during the stream. Describe the strongest boundary the implementation actually provides. A supported fence, ordering token or serialized release protocol may improve that boundary, but each needs its own semantics and recovery evidence. A timestamp alone is not an atomic fence.
Choose a coarser, explainable contract if the architecture cannot support the desired one. For example, a protected workflow can hold an entire answer until eligibility is resolved rather than immediately streaming excerpts. This trades responsiveness for a simpler release decision, though it still requires correct handling of the check-to-release interval. The trade-off belongs in product requirements and operational acceptance. Do not market a stringent boundary while implementing the easiest admission-time check and hoping the interval is too short to notice.
The timing view separates policy change from an answer's progress. It is not a claim of atomic coordination. The remaining check-to-release interval is the part an implementation must explain rather than hide behind a final-check label.
Compare the options by the promised boundary
The options are not mutually exclusive products. Synchronized metadata can narrow candidate retrieval, current checks can verify eligibility and release coordination can govern pending answers. Adding all three without a contract can produce expense and complexity while leaving the same gap. Select the minimum combination that supports the actual source semantics and disclosure obligation, then verify its behavior under delayed synchronization, policy-service failure and already-running work. A design is defensible because its evidence matches its promise, not because it contains the largest number of controls.
For a public documentation assistant, per-person revocation may not be the dominant requirement. Content publication and removal still matter, but copying an enterprise ACL model onto public pages can create unnecessary dependencies. For a protected project assistant, source membership and derived answers may need joint treatment. For regulated or specially restricted content, the accepted boundary may require withholding output whenever present eligibility is unresolved. These examples are proposed decision contexts, not legal categories or recommendations to assign a particular policy without review.
Use an option record with six fields: source authority, supported relationships, policy freshness evidence, derived-content disposition, pending-response treatment and unavailable behavior. Name what each option does not establish. A synchronized ACL proves the indexed decision used at query time; it does not by itself govern an old report. A current check proves the authority's decision under that check's semantics; it does not retroactively erase delivered bytes. A release fence governs its supported boundary; it does not repair omitted endpoints.
Reject a selection when its prerequisites cannot be observed. If the source cannot identify relevant revisions, a proposed revision-based control may not be implementable. If an index cannot expose applied permission state, the team may need independent fixtures and a conservative serving block instead of a reassuring dashboard timestamp. If a cache cannot identify contributing sources, it may need to be excluded from protected reuse. Choosing not to cache a source class is a legitimate architecture decision, not a failure to optimize.
Treat derived answers as protected objects
An answer made from several documents is a new serving object with dependencies. Its access requirement cannot safely be inferred from the least restricted source. Record which protected revisions contributed to it and what rule determines its audience. If contribution cannot be established reliably, the application needs a conservative disposition such as refusing reuse or recomputing with currently eligible context. This proposed approach values explainable scope over a cache hit whose provenance was discarded for convenience.
Separate answer identity from semantic similarity. A similar question is not proof that two users can receive the same response. Establish tenant and access eligibility before using similarity to select a reusable artifact. An answer prepared under a manager's broad scope must not become a general response simply because its wording looks harmless. Titles, numbers and omissions can reveal information without a direct quotation. Text redaction after retrieval is not a substitute for choosing eligible input and governing the resulting artifact.
Amazon Bedrock's query configuration guidance documents metadata filtering of retrieved content, with requirements depending on the knowledge-base configuration. That is a retrieval capability. This paper's inference is narrower than a security claim: the surrounding application must decide whether its metadata and filter construction represent the reader's trusted policy, and must govern caches or history that do not repeat retrieval. A filter parameter should not be presented as an automatic end-to-end revocation protocol.
Specify a reusable-answer record that serves the decision. Useful proposed fields include tenant, answer identifier, source revision references, policy observation identity, task configuration and release disposition. Keep sensitive body text out of broad audit logs. The record should let an owner locate affected answers and explain why they were withheld or reused. It should not become a second unrestricted search corpus. Apply its own access and retention controls, since source associations can themselves reveal confidential projects or relationships.
Handle history, citations and exports independently
Conversation history has a different user expectation from a new answer. A reader may expect to see an earlier discussion, but the server may now hold content they are no longer entitled to receive. Decide whether history is reauthorized on access, whether affected messages are replaced with a restricted-content notice, or whether a separately approved historical-view policy applies. The choice needs a source-owner decision and a test. Do not quietly inherit whichever behavior the conversation database offers by default.
Mixed-source messages require careful treatment. If only one excerpt is affected, selectively hiding it may be possible where provenance and rendering boundaries are reliable. If the message synthesizes inseparable protected facts, withholding the whole message can be more defensible. Rewriting it with a model creates a new processing and disclosure decision; the rewrite needs eligible context and verification too. A message saying “something was removed” can leak a restricted topic if its title or notification is not considered part of the serving surface.
Review citation and export routes by direct access. An answer may correctly omit protected text while a citation label reveals a document name. A signed file URL may remain usable under its own lifetime even after the application changes permission. Record who issues such URLs, what authority they encode and whether the revocation contract can actually control them. If it cannot, state that limit and choose a different serving mechanism for sources whose requirement is stricter. Obscurity of an identifier is not object authorization.
Elastic's security limitations note that filtered aliases are not a secure document-access mechanism and describe aggregate-information limitations even with document-level security. Review the actual APIs exposed by an assistant, including counts and metadata, rather than assuming excluded result bodies imply no information can escape. This is a source-specific limitation to assess, not a claim that every search backend leaks the same fields or that an ordinary application necessarily exposes those APIs.
Make in-flight and streaming behavior explicit
Pending work should carry enough identity to be affected by a permission change. A generation job can retain the requesting subject, tenant, contributing source revisions and its current release state without logging every passage. Define how an owner locates jobs affected by a scope change and what cancellation means. A request to stop generation does not prove that a queued notification, report export or retry will also stop. Each downstream disclosure needs a disposition under the same approved contract.
Streaming changes the boundary from one held response to a sequence of deliveries. Before enabling protected streaming, decide whether eligibility is checked at admission, before the first byte, at supported intervals or under a coordinated session policy. State the accepted effect of a change during the stream. Bytes already sent remain delivered even if the UI clears them. A visual replacement is not evidence of nondisclosure. Where the application cannot meet the required contract with streaming, hold the answer or limit streaming to a separately classified source path.
Use distinct states for denied, unavailable and changed-scope results. Denied means the trusted authority does not permit the relevant disclosure. Unavailable means the authority or necessary evidence could not be obtained. Changed scope means pending work was assembled under an observation no longer usable for the required boundary. These states may share a user-safe message, but operators need the distinction to recover correctly. Automatically retrying every withheld answer can recreate the same uncertainty and increase exposure or cost.
If recomputation is allowed, start from eligible context rather than asking the model to forget the removed passage within an existing prompt. Bind the new attempt to a new evidence record and retain only the minimal relationship to the withheld attempt. Check the output using the accepted fixtures before release. A regenerated answer can still reveal protected facts through cached tool results or conversation context, so recomputation is not inherently safe. Its inputs and serving path need the same scope review as a fresh request.
Recover from missed changes without reopening access
Permission events are useful acceleration signals, but the design should not assume each consumer receives them once and in order. Record the event or revision identity, the serving stores expected to apply it and the observation used to close each application step. A duplicate should not reverse a restriction, and an older update should not replace a newer decision. Where the source lacks a comparable revision, design an explicit reconciliation method rather than inventing an order from local receipt timestamps.
A proposed reconciliation job can compare the authority's supported current state with serving metadata for a bounded source family. It should report mismatches, age and unobserved relationships, not silently declare every missing record revoked. Confirm the scope before broad repair. A connector outage can mean “not observed,” not “permission removed.” Conversely, a successful ingestion run can leave an inherited ACL unchanged if the connector does not support that change path. The reconciliation must reflect actual source semantics and independently verify the repair it applies.
Restoration deserves its own check. A recovered index or cache can contain permissions that were valid when backed up but are no longer valid now. Keep restored protected data out of normal serving until the required policy reconciliation is complete. Test this state rather than relying on a backup timestamp alone. The related enterprise AI deletion paper discusses obsolete derived copies; revocation differs because the data can remain valid for other readers even while one relationship must stay removed.
Recovery actions need narrow authority and durable progress. Name which source family is paused, which revisions are reconciled and which observations permit reopening. Do not grant broad reader access simply to make a repair script easier. If the repair affects caches, history and previews, close each path independently. An unobservable path remains held or explicitly outside the accepted scope. This can reduce availability for a protected feature, but treating unknown state as permission would preserve a different end state from the one the design promised.
This ledger compares independent surfaces rather than presenting them as sequential steps. Its hold actions are proposed dispositions to adapt to an approved contract, not provider-managed guarantees.
Measure the accepted boundary and its operating cost
Measure from the declared authoritative event to the observations required by the contract. Keep index application, cache disposition and response withholding separate. Report the sampling interval and uncertainty: a probe every thirty seconds cannot prove a millisecond-scale transition. Preserve failed attempts and mixed results. If one endpoint denies while another discloses, a single averaged delay hides the important finding. The owner needs the unresolved surface and its exposure condition, not just a reassuring mean.
Include positive controls in every exercise. A retained reader should still retrieve the valid source, and an unrelated public source should remain usable under its own policy. These controls distinguish correct scope from a broken system that denies everything. Test both fresh and existing sessions, warmed and empty caches, direct references and ordinary questions. Choose representative fixtures for supported relationships, including inheritance or temporary grants where actually used. A never-authorized identity is valuable, but it does not replace the formerly authorized reader who carries stale state.
Operational and security consequences meet at the unavailable path. Current checks can increase latency and dependency load. Blocking a source family can interrupt useful work. A broad invalidation can reduce reuse and trigger expensive recomputation. Price those consequences using measured request families and accepted-output counts, then compare them with the chosen confidentiality obligation. The AI workflow economics paper provides the cost-accounting frame. Do not lower the promised boundary silently to improve a performance graph.
Record an error budget for service behavior without making it an authorization allowance. A product can accept some unavailable protected responses while still forbidding known unauthorized disclosures. Those are different metrics with different owners. Avoid a composite “quality” percentage that trades a serious scope failure against many fast answers. Report denied, unresolved, recomputed and accepted outcomes separately. Review the practical user messaging so a temporary policy dependency failure does not falsely tell someone they permanently lost access.
Decision checklist and limits of the recommendation
Use this review checklist before approving a protected source class. The application owner should identify the actual serving paths and the security reviewer should confirm the permitted boundary. Require evidence rather than a feature screenshot. The checklist is a proposed acceptance artifact, not a security certification or an exhaustive threat model. It should be supplemented by review of the deployed identity model, source semantics, provider configuration and application attack surface.
- [ ] The subject-resource relationship and source authority are named.
- [ ] The effective event and supported policy-freshness observation are explicit.
- [ ] Retrieval, caches, pending responses, history and direct exports have separate dispositions.
- [ ] Derived answers retain sufficient provenance or are excluded from protected reuse.
- [ ] The check-to-release interval and streaming limits are stated without an atomicity claim.
- [ ] Missed, duplicate and older updates cannot silently reopen serving.
- [ ] Restored copies remain restricted until current policy reconciliation is accepted.
- [ ] Denied and unavailable states are distinguishable in restricted operational evidence.
- [ ] Formerly authorized and retained-access controls have observed results.
- [ ] Uncovered source relationships and serving paths remain visible in the release decision.
This recommendation does not apply unchanged to a public-only assistant, where publication state may be the primary boundary, or to a source whose permission semantics cannot be represented by the application. It also cannot guarantee confidentiality from a compromised privileged operator, remove previously delivered external copies or establish compliance with a particular law. Those questions require additional controls and qualified review. State them plainly rather than expanding the word “revocation” until it appears to solve every information-governance problem.
The next step is to complete one option record and run the revocation test playbook against synthetic fixtures. Bring the serving-surface inventory, policy observations, withheld-answer evidence and unresolved intervals to an AI workflow review. The useful outcome is a narrower, defensible disclosure promise and a design that can demonstrate it. A diagram, a provider feature or this paper alone is not that evidence.