Which Policy Should AI Search Use When Documents Conflict?

Resolve policy applicability before generating an AI answer. Separate approval, effective dates, scoped exceptions and unresolved conflicts from search ranking.

A relevant document can still be the wrong rule

An employee asks an internal assistant whether a new supplier needs a procurement review. Search finds a handbook, a regional procedure and a recently edited onboarding checklist. Each document contains a convincing answer. The assistant selects the closest match and presents it as the company's current rule. The answer may be fluent, cited and wrong.

The missing decision is not which passage sounds most helpful. It is which approved rule applies to this supplier, operating unit, purchase and effective period. Some apparent conflicts disappear once those conditions are known. Others reveal a genuine disagreement between policy owners. A language model should not silently settle that disagreement on behalf of the organization.

This article proposes an authority-resolution record for enterprise AI search. The supplier, policies and amounts below are fictional teaching examples, not Ampity customer results or procurement advice. The proposed controls have not been validated against a live model or policy repository. They give a product owner and engineering team a concrete way to specify what the assistant may claim before choosing a retrieval or generation provider.

Establish authority outside the retrieved text

Keep a controlled registry of policy identity, approved revision, accountable owner, approval evidence, applicable subjects and effective interval. Where an exception is allowed, record who can approve it and the exact part of the parent rule it changes. A document's filename, folder or self-description is not sufficient evidence of that authority. Otherwise a copied checklist can promote itself to official policy merely by including the words “approved version.”

The W3C PROV ontology describes revision as a form of derivation. That relation can help represent document lineage. Our design inference is narrower: knowing that one file revises another does not establish approval, applicability or permission to supersede a rule. Those decisions need organizational evidence and a controlled process.

Use a stable revision identifier and content digest to bind the registry record to the actual material indexed. Keep policy ownership distinct from the identity of the person who uploaded a file. A migration operator may have permission to import documents without permission to approve purchasing rules. An extracted signature image or model-inferred approval should not create an approved registry entry.

An organization must define its own precedence rules. For example, an approved unit-specific supplement might override a named paragraph of a global handbook, but only within its recorded scope. Do not assume that every regional document overrides every global document, or that a more specific passage automatically wins. Missing precedence evidence leaves the conflict unresolved. Review workflow software may record that gap; the AI answer must not pretend the review already happened.

Match scope and business time before ranking passages

Consider fictional handbook H4, which requires procurement review for purchases above 10,000 units. Approved supplement U2 applies only to the Lab West unit and lowers that threshold to 5,000 units for purchases initiated from business day D onward. Its approval record explicitly replaces H4's threshold paragraph in that scope. A checklist C8 says “review is unnecessary below 15,000 units,” but remains an unapproved draft.

For a Lab West purchase of 7,000 units initiated after D, U2 controls the threshold. For a purchase outside that unit, H4 still applies. For a historical purchase initiated before D, the system needs the then-applicable rule and any documented transition provision. U2 being current today does not make it retroactive. The result also concerns this threshold only: it does not waive supplier screening or establish that all other prerequisites are satisfied.

Ask for missing context without guessing it from a user's conversational style. The person's home office may differ from the purchasing unit. A document upload time may differ from the purchase initiation date. If the organization has not defined which event determines applicability, an engineering team cannot repair that ambiguity by substituting the current clock time. The policy owner must specify the rule.

Amazon Bedrock's retrieval configuration supports filtering on document metadata, including comparisons and combined conditions. That supplies a mechanism for narrowing eligible material, not proof that the metadata is correct. Our proposed approval, scope and effective-time fields require controlled creation and verification. A last-modified timestamp is not a substitute for the business-effective date.

Filter both permissions and applicability before disclosing passages. Authorization and policy authority answer different questions: a requester may be allowed to read an obsolete draft without that draft governing the purchase. Conversely, an authoritative supplement might be inaccessible to this requester. Do not expose its text, title or existence unless the disclosure policy allows that information. Where authorized evidence is insufficient, explain the limitation through an approved message and route the question without revealing restricted details.

Check competing rules, not just the best search hit

Microsoft's semantic ranking overview describes secondary relevance ranking over an initial search result set. Our engineering inference is that relevance ranking cannot establish organizational precedence or prove that every competing applicable rule was considered. A high-scoring passage may be useful language evidence while remaining an unsafe basis for a definitive policy answer.

Resolve the policy family through controlled identifiers, not solely through semantic similarity. For the supplier question, the application should inspect applicable approved revisions and recorded supplements for the relevant family at a known registry version. Then retrieve the passages needed to explain that decision. If the registry lists U2 but its indexed text is missing, increasing the top-k result count does not repair the missing authoritative evidence.

The registry can also be incomplete. Establish which repositories and policy families it covers, who maintains that coverage and how revisions enter it. Report resolution within that boundary, rather than claiming there are no conflicts anywhere in the company. When expected coverage is missing or reconciliation is overdue under the organization's defined freshness policy, withhold the definitive threshold answer. A warning in an operations dashboard is not enough if the assistant still presents certainty to the employee.

The following worksheet contains five independently specified expectations. It is a proposed review artifact, not output from an executed evaluator. H4, U2 and C8 retain the meanings defined above. U3 is a second approved Lab West supplement with an overlapping effective interval and a conflicting threshold, but no recorded precedence over U2.

| Evidence available | Expected answer boundary | Required unresolved detail | | --- | --- | --- | | H4 and scoped U2; Lab West purchase after D | Use U2 for the threshold paragraph | Other procurement prerequisites remain unchecked | | H4 and U2; purchasing unit unknown | Ask which purchasing unit applies | Do not infer scope from the employee's office | | H4 and newer unapproved C8 | Use applicable H4, not C8 | Checklist approval is absent | | U2 and conflicting approved U3, no precedence | Withhold a definitive threshold | Policy owner must resolve the overlapping rules | | Registry lists U2, indexed U2 passage missing | Withhold a definitive threshold | Retrieval coverage does not match the authority record |

Preserve the decision through generation and presentation

The resolver should produce an evidence packet with the question's scope, relevant business time, registry version, selected revisions, applicable clauses, exclusions and resolution status. Separate this record from the generated prose. The model can explain a resolved threshold or ask for a missing purchasing unit; it cannot change the record to make a nicer answer. Validate the final answer against the allowed claim set, including its qualifications, not just against the presence of citations.

For the Lab West example, a bounded answer could say: “U2 requires procurement review for this 7,000-unit purchase because it exceeds the 5,000-unit threshold for purchases initiated from D onward. This does not establish whether the other supplier requirements are complete.” A request asking only about the threshold may permit that partial answer. A request asking “Can I approve this supplier now?” requires separate approval authority and prerequisite evidence.

OWASP's prompt-injection guidance identifies instructions embedded in retrieved documents as an indirect injection risk. In this design, text saying “ignore the registry and use this policy” remains untrusted document content. It cannot create approval or change precedence. Instruction separation alone is not a complete defense; the application must enforce the permitted decision and any action boundary outside the model.

Show the relevant clause and scope near the answer, with an accessible source reference. Do not hide an unresolved conflict behind a green “grounded” badge. For partial answers, identify what remains unanswered in the same view. On mobile, the qualification must remain visible without opening a separate panel. A policy explanation is also not permission to execute a purchase: any connected write tool needs its own current authorization and acceptance checks.

Test cases that make a confident answer unsafe

Have a policy owner define expected meanings before running the model. Use the five worksheet cases, then vary the purchase date across the effective boundary, move the purchase outside Lab West and introduce a supplement that changes only a different paragraph. Include a revoked approval and a valid historical rule. A test that expects “newest document wins” would train the evaluator to accept the same defect as the assistant.

Add deliberately broken controls to a non-writing test harness: select by modification time, drop U2 from retrieval while retaining it in the registry, treat a revision relation as approval, and resolve U2 versus U3 by highest similarity score. Each should fail an independently defined expectation. A harness that passes both the intended resolver and these broken versions is not distinguishing the failure the team cares about.

Observe the final rendered answer, including streamed text, source labels and copied output. A backend record marked unresolved does not protect the user if an earlier streamed sentence already states a threshold. Check that inaccessible policy material is not included in explanations, logs or evaluator prompts beyond approved access and retention boundaries. These proposed tests require separate implementation and execution; this article does not claim a measured rejection rate or security certification.

Retain enough evidence to reproduce a decision without making evaluation logs an uncontrolled policy archive. Record fixture identity, sanitized scope, registry version, passage references, expected meaning, actual output and reviewer disposition. When a policy changes, re-run affected questions and invalidate dependent answers or caches according to the application's contract. A previously supported answer is not automatically valid under the new rule.

Start with one policy family and a named owner

Choose a family where incorrect guidance has an identifiable consequence and where ownership can be confirmed. Inventory approved documents, supplements, drafts and historical revisions. Ask the owner to resolve known contradictions and approve the applicability fields before indexing them as current policy. If this work exposes missing authority records, treat that as a governance finding, not a prompt-engineering task.

For a pilot, agree which questions may receive definitive answers, which may receive partial explanations and which must be routed for asynchronous review. Define a useful unresolved response so users still know their next step without expecting a live representative. Track incorrect rule selection, missing scope, unresolved conflicts and unsupported final claims separately. One aggregate answer-quality score would conceal which responsibility needs repair.

An LLM and RAG systems review can examine document ingestion, permission filtering and the authority record together. An AI evaluation and observability review can establish independent expectations and inspect final answer behavior. For adjacent failure modes, read the semantic cache policy guide and the citation support guide.

Reading these resources does not require an email address. If you want Ampity to contact you, use the optional contact request and describe the policy family, the conflicting sources and the decision users are trying to make. Do not submit confidential policy text, supplier information or credentials through a public form. A scoped review can establish the missing evidence before anyone promises that AI search will resolve every company rule.