A Citation Does Not Prove That an AI Answer Is Supported
Check AI answers claim by claim. Use a worked citation worksheet to distinguish source pointers, actual support, conflicting evidence and permission to act.
Check what the source supports, not whether a link appears
An AI answer can point to a real document and still reach the wrong conclusion. A source may support one clause but not another, describe a different customer, or apply only after an approval that has not happened. Counting citations will not expose these mistakes. Test the important assertions against the exact evidence and its applicable conditions.
This matters when an employee uses an assistant to answer a customer, interpret an operating policy or prepare a business change. A polished paragraph with three references can look safer than an uncited paragraph even when both contain the same unsupported promise. The interface should make evidence easier to inspect, not turn the presence of a reference into an endorsement.
This article proposes a review method, not a guarantee about any model. The refund policy, amounts and responses below are invented teaching examples. They are not Ampity customer results, a real contractual policy or legal advice. The useful output is a reusable assertion worksheet that your product owner and engineering team can adapt to their own answer contract.
Separate citation integrity from answer acceptance
Three different checks are often compressed into the word “grounded.” First, a citation needs to resolve to the intended source passage. Second, that passage needs to support the assertion made. Third, the source needs to be applicable to this requester, subject and moment. Passing the first check does not settle the other two.
Anthropic documents that its citations feature produces valid pointers into the supplied documents and returns source-location information. That is valuable for verification. Our engineering inference is that the application must still decide whether the cited material supports the whole business assertion and applies to the current task. A valid pointer is not an application-level acceptance decision. See Claude's citation documentation.
Google's grounding documentation describes checking an answer candidate against supplied facts, including claim-to-evidence associations. It also distinguishes full support from partial support: getting one part of a claim right does not establish the whole claim. This makes claim-level inspection a useful complement to citation rendering. See Google Cloud's grounding checks.
Neither citation presentation nor a support score establishes that the selected evidence set contains the latest applicable policy. Source selection, permissions and revision validity remain separate responsibilities. An accurately supported statement from the wrong source can still be the wrong answer for the user.
Work through a response that looks reassuring but fails
Consider a synthetic cancellation assistant. The requester asks: “The venue cancelled our event. Can I tell customers that everyone will receive a full refund today?” The evidence package contains three deliberately different records.
Policy R7 permits a full refund of the registration fee after an operations manager approves a venue-cancellation refund batch. Payment processing fees are excluded. A payment runbook says an accepted refund request may remain pending while the payment provider processes it. The event record confirms the cancellation but shows the refund batch as awaiting approval.
A candidate answer says: “Yes. Everyone will receive a full refund today because the event was cancelled. The cancellation policy permits refunds, and the payment runbook describes the process.” Both cited documents exist. Both concern refunds. Yet the answer substitutes permission for completion, omits the excluded fee, assumes approval and invents a deadline.
The reviewer should reject the promise, not merely ask the assistant to add another citation. A safer answer separates the established cancellation from the unapproved refund batch, explains the policy's fee boundary and says that completion timing is not established by the supplied records. It can offer to locate the approval status without claiming that approval or payment has occurred.
Use an assertion worksheet before discussing a score
Split statements at the point where a different piece of evidence or condition is needed. Do not rely on sentence punctuation alone. “Everyone will receive a full refund today” contains a population claim, an amount claim, a completion claim and a timing claim. Reviewers should be able to disagree with one part without accepting or rejecting every other part blindly.
| Assertion | Evidence check | Decision for this example | | --- | --- | --- | | The event was cancelled | Event record confirms the same event | Supported | | Every customer is eligible | Applicable registrations and exclusions not checked | Not established | | Refund includes every fee | R7 excludes processing fees | Contradicted | | Refund is approved | Batch remains awaiting approval | Contradicted | | Funds arrive today | Runbook allows pending processing | Not established |
The table is an adjudication aid, not a mathematical model. “Not established” means the inspected evidence does not justify the assertion; it does not prove the assertion false. “Contradicted” means relevant evidence opposes it. Preserve that distinction in the evaluation record and in user-facing wording.
Add the exact source revision, passage locator, applicable entity and reviewer reason to the underlying record. A screenshot of a green score is not enough to reproduce the decision. Keep access controls and retention limits on this record, because it can contain customer context or restricted policy text. Evaluation evidence must not become a new uncontrolled copy of business information.
Keep conditions attached when retrieving and summarizing
A retrieval chunk that contains the permitted action but omits its prerequisites invites an overbroad answer. Include enough surrounding context to retain approval requirements, exclusions, effective dates and the subject the rule governs. More context is not automatically better: unrelated material can make it harder to identify which rule applies.
Test the same question with the approval paragraph missing, a superseded policy present and two similarly named events in the evidence package. Expected behavior should change with the available evidence. When a required condition cannot be determined, the assistant should identify the missing fact rather than complete the answer from a familiar pattern.
Separate “the document says” from “the current event qualifies.” A policy is evidence about a rule. The event record supplies facts needed to apply that rule. A derived eligibility conclusion requires both, and the application should retain that reasoning boundary even when the interface presents a short answer.
If sources conflict, do not let an undocumented model preference decide which wins. Define an authority rule appropriate to the domain, such as an approved policy revision taking precedence over an informal internal note. If the rule cannot resolve the conflict, report the uncertainty and identify the responsible owner. A newer timestamp alone need not imply higher authority.
Measure failures that matter to the decision
Start with a small, adjudicated set covering ordinary answers, excluded cases, conflicting revisions, missing prerequisites and plausible unsupported numbers. Record which assertions matter most to the user's next step. Missing a decorative fact and inventing an eligibility promise should not be treated as interchangeable defects.
Keep citation resolution, assertion support, evidence applicability and answer usefulness as separate review dimensions. An answer can cite accurately but fail to answer the question. It can also answer the question using evidence the requester should not see. Improving one dimension must not hide a regression in another.
If you use an automated evaluator, compare its judgments with reviewer-adjudicated examples before trusting its threshold. Include disagreement cases, especially partial support and conflicting evidence. Do not present a model-produced support score as a calibrated probability that a customer-facing promise is correct. Decide how much automation is appropriate from the consequence of an error, not from the convenience of a numeric output.
When generation, retrieval or source-publication behavior changes, rerun the relevant examples and examine changed assertions. A release can improve average answer fluency while damaging an exception the business relies on. Keep the observed failure, evidence package and accepted response together so an engineer can reproduce it without guessing which documents were retrieved.
Design the interface to expose uncertainty without blocking help
Put citations next to the assertions they support. A single source list at the bottom of a long response makes it difficult to tell which reference supports which conclusion. Where practical, let the reader inspect the relevant passage and source revision without losing their question or conversation state.
Use plain distinctions such as “the cancellation is confirmed” and “refund approval is not confirmed.” Avoid turning every incomplete response into an empty refusal. The assistant can still explain the rule, show the established facts and ask for one missing detail. It should not ask for unnecessary personal information just to reveal an answer.
For a service website, a useful assistant can answer from published material and offer an optional enquiry when the evidence runs out. It should not pretend that a representative is online or that an answer constitutes a commercial commitment. For an internal workflow, source-supported explanation also remains separate from permission to change a record or send a message.
Take one real question through the full acceptance path
Choose a question your users actually ask and collect the permitted source revisions needed to answer it. Write the expected assertions and the facts that would make each assertion inapplicable. Then produce candidate answers with one condition missing, one source contradicted and one unsupported promise inserted. Review whether your current tests catch those specific defects.
The next action is not “add citations everywhere.” It is to identify one unsupported assertion your present process would accept, decide how the product should respond and add a reproducible regression case. That exercise reveals whether the missing control belongs in retrieval, source governance, evaluation or the interface.
For the wider release decision, read governing changes to production AI systems. For retrieval-specific regression work, see why a retrieval change can break AI answers. These address adjacent decisions rather than treating citation display as the entire reliability problem.
If your team needs help defining an answer contract, Ampity's AI engineering work can start from the question, permitted evidence and consequence of a wrong answer. You can use this worksheet independently. Contact is an optional next step, not a condition for applying the method.