When a Semantic Cache Returns an Outdated Policy
Treat a semantic cache hit as a candidate, not current policy evidence. Check source revisions, audience scope and invalidation before reuse.
A similar question does not prove the answer is still valid
A semantic cache can return an answer created for a similar earlier question. Before serving it as current policy, check that the source revision, audience scope and task context remain valid. Similarity is a retrieval signal. It is not evidence that a policy is unchanged or that today's requester may see yesterday's answer.
This article proposes an answer-reuse contract for engineering teams. Its policy example is hypothetical, not an account of an Ampity engagement. The recommendation concerns cached generated responses. Provider prompt caching, which reuses computation for input prefixes, is a different mechanism and should not be described as serving a previously generated business answer.
A fast wrong answer is especially difficult to notice when it retains confident wording and plausible citations. The application may bypass retrieval entirely on a hit, so a correct source update does not necessarily affect the returned text. Review the reuse path as a separate information-delivery path, not an invisible performance option.
Separate similarity matching from eligibility to reuse
Microsoft's semantic-cache lookup reference describes vector-proximity matching and cache partitioning. It also warns that similar prompts can yield incorrect or outdated cached responses. The controls proposed here add application-specific eligibility checks; they are not a claim that a cache gateway understands your business policy.
First determine which cache partition the current request may search. Then find candidate entries and check whether each candidate remains eligible. A close semantic match across different organizations, policy editions or effective dates should not become a shared answer merely because its distance score is favorable.
Keep those decisions observable. Record a cache candidate, a validation outcome and a reason for reuse or bypass. A single “hit” counter can combine legitimate reuse with rejected candidates. It cannot explain whether the optimization preserved the information boundary.
Do not infer freshness from the embedding. An embedding represents content under a chosen model; it does not automatically encode a subsequent source change. Increasing the similarity requirement may reduce unrelated matches, but it does not make an old answer current. Treat matching and freshness as different tests.
Bind the answer to the evidence that produced it
For a cacheable answer, retain the source identifiers and revisions used, the relevant effective date, the permitted audience and the task-contract version. Keep the generation configuration separately where useful for investigation. A model name alone does not identify the policy basis of the answer.
Prefer a clear evidence snapshot over a timestamp with no meaning. “Generated at 10:00” says when the text was produced, not whether the source was authoritative at 10:00. If several sources support the response, the eligibility check must cover the relevant set rather than only the first cited document.
Define what happens when the evidence cannot be verified. An application might regenerate against current authorized sources, return a scoped unavailable message or defer the decision. It should not silently treat a missing source record as permission to reuse the answer indefinitely.
Some tasks depend on more than document revisions. A policy may differ by jurisdiction, account plan, event date or contractual exception. Include the decisive context in the reuse rule. If the system cannot reliably identify those dimensions, the appropriate answer may be to exclude that task family from semantic caching.
Worked example: the reimbursement period changed
Imagine a synthetic internal assistant answering questions about a reimbursement policy. Revision A permits submission within 60 days. Revision B changes the period to 30 days for claims arising after a specified effective date. An answer generated from revision A remains in the response cache.
The next requester asks a paraphrase of the same question. A semantic match can be excellent while the correct response depends on which revision applies to the claim date. Expiring the entry after an arbitrary hour does not prevent it being served during the hour immediately after revision B becomes effective.
| Current condition | Proposed reuse decision | Evidence required | | --- | --- | --- | | Same authoritative revision and applicable claim date | Candidate may be reused | Current eligibility and audience checks pass | | New revision applies to the claim | Reject cached answer | Rebuild using the applicable source revision | | Earlier revision still governs a historical claim | Use only a historical-policy path | Verified date and explicit historical framing | | Claim date is missing | Ask for necessary context | Do not select a policy edition by similarity | | Source authority cannot be confirmed | Bypass or defer | Approved unavailable-information behavior |
The example illustrates why “latest policy” and “policy that governed this event” are not always the same question. A system that always switches to the newest text can also be wrong for a historical request. Write the applicable-version rule before choosing the cache key.
Invalidation needs a publication boundary
Identify who can publish a new policy revision and how that publication changes answer eligibility. Possible designs include advancing a revision namespace, invalidating dependent entries or checking a current revision manifest at read time. Choose based on the freshness requirement and the ability to observe failed updates.
Deleting entries is useful, but deletion alone can race with generation. A request started under revision A may finish after the invalidation and repopulate the cache with an A-based answer. Before storing its result, compare the generation's evidence snapshot with the current publication state. Reject an obsolete fill rather than restoring what the invalidation removed.
The cache-aside pattern reference discusses consistency and invalidation trade-offs. Applying those ideas to generated policy responses requires an explicit dependency record. Unlike a simple cached database row, one answer may summarize multiple records and a task-specific interpretation.
Give invalidation delivery an owner and observable status. An event saying “policy updated” is not proof that every region, local process and cache namespace applied the change. If the application requires strict current-policy delivery, use a verified read-time eligibility gate or bypass reuse while publication state is uncertain.
Audience partitioning is necessary but not sufficient
Use an authenticated scope, not a user-supplied tenant string, to select the candidate partition. A cache entry must not acquire wider visibility just because it is stored in a shared service. Keep access rules consistent with the underlying source and avoid unnecessary sensitive material in cached text.
A partition can still contain stale permissions. Group membership or document access may change after generation. An entry originally valid for a requester can become inappropriate later even when the policy text stays unchanged. Recheck the relevant current permission boundary before delivery; do not use the old generation-time permission snapshot as continuing authorization.
If a cached answer blends public and restricted material, the combined result needs the restrictions of the material it exposes. Removing the visible citation does not remove the sensitive information from the prose. Treat derived answers as information-bearing artifacts, not harmless metadata.
These controls do not prove that response caching is suitable for every sensitive workflow. A current-source read or a deliberately noncacheable response may be simpler to operate and easier to verify. Explain that trade-off instead of presenting partitioning as a universal safety guarantee.
Test stale hits, not only expected cache hits
Create a fixture where an answer is generated, the source changes and a paraphrased question arrives. Inspect the actual text returned and the evidence revision, not just an invalidation message. Repeat with an in-flight generation crossing the publication boundary and with a local cache that misses the update event.
Include a changed effective date, a removed document, a revoked requester scope and an ambiguous historical question. The expected result may be bypass, clarification or refusal to provide unsupported policy advice. Decide those outcomes before the test so a cached answer cannot pass merely because it sounds plausible.
Measure validated reuse separately from candidate hits. Track stale-candidate rejection, unexpected cross-scope candidates and regeneration load. Avoid copying entire policy answers into general telemetry when identifiers and controlled investigation access suffice. Retain enough evidence to explain a failure without creating another unrestricted response store.
Limitations and a practical starting point
This proposed contract cannot make missing evidence authoritative or repair an inconsistent source system. It also adds metadata and read-time work that can reduce the apparent benefit of caching. Measure that cost together with the consequences of serving stale information. A low hit rate may justify removing the response cache rather than weakening eligibility rules.
Start with one low-risk, clearly versioned question family. Write its permitted audience, applicable-version rule, invalidation owner and bypass behavior. Run the stale-hit fixtures before increasing scope. Use the retrieval-change article to examine the fresh-generation path and the AI release-gate playbook to organize acceptance evidence.
For the economic comparison, read cost per completed AI task. If you need help implementing evidence and access controls, explore production AI systems or ask a workflow question. Resource access remains independent of submitting contact information.