Missing Evidence Is Not an Extraction Error

Distinguish absent, unreadable and contradictory document evidence from extraction failures. Route each state to an owner and a useful next step.

Ask whether the evidence exists before retrying extraction

When a required value is missing from an AI document-intake result, first distinguish absent evidence from unreadable evidence, contradictory evidence and a processing failure. Retrying extraction can help with some processing failures. It cannot supply a page that was never submitted or decide which of two conflicting approved records is authoritative.

This article proposes an exception-routing model for engineering and intake teams. Its invoice and application examples are synthetic. They do not establish customer results, legal evidence standards or permission to execute financial changes. The practical question is narrower: what information would let the assigned person resolve this item safely?

A generic “AI failed” queue obscures that question. Reviewers open the same document repeatedly, rerun the same job and eventually type a plausible value just to move the item along. The application then records a completed field without showing whether it came from the document, another approved source or an unsupported guess. Separate those states before optimizing throughput.

Give each exception a reason and an owner

Use reason codes that describe the available evidence, not just which service returned an empty field. An empty candidate is an observation. “Required page absent” is a conclusion that needs inspection under a defined intake rule. Keep the observation and the review decision separately so later investigators can see how the conclusion was reached.

| Observed condition | Question to resolve | Proposed next step | Accountable owner | | --- | --- | --- | --- | | Submitted document lacks the required section | Was the submission incomplete? | Request the relevant source or approved alternative | Intake coordinator | | Section exists but is unreadable | Would a better source reveal the value? | Request a clearer scan or original file | Submitter through intake coordinator | | Two relevant sources disagree | Which source governs this decision? | Hold the effect and resolve authority | Business decision owner | | Processing request failed | Was any result produced or stored? | Inspect job state, then retry within policy | Processing operator | | Value extracted but interpretation ambiguous | Which interpretation is supported? | Review context or ask for clarification | Qualified reviewer | | Candidate conflicts with a current record | Is the source stale or the record wrong? | Compare revisions and follow the correction process | Record owner |

These owner labels are roles to assign, not a claim that every team has separate staff for them. One person may cover several roles, but the queue should still make the responsibility explicit. A technical operator should not acquire business authority simply because they can restart the extraction job.

Preserve evidence without pretending it proves completeness

Store the exact document revision a candidate came from, along with the relevant page and evidence anchor. If a reviewer requests a replacement, preserve the relationship between the old and new submissions. Do not overwrite the original while leaving a prior decision attached to what now appears to be the same file.

Amazon Textract's Block API reference describes page, geometry and relationship data associated with detected elements. Those structures can support source navigation. They do not, by themselves, prove that an expected field was absent from the entire submission.

Completeness depends on a separately defined requirement. An application might require a signature page, a referenced attachment or a clearly identified account statement. Write that requirement in the intake rule and specify which document types it applies to. A detector not returning a signature candidate is not automatically proof that the signature does not exist.

If the application allows evidence from another source, identify that source and the authority for using it. For example, a reviewer may consult an approved master record under a defined procedure. The resolved value should retain that provenance rather than being represented as if extracted from the submitted file. This distinction becomes important when someone later asks what the submitter actually provided.

Reprocessing and requesting evidence are different workflows

Reprocessing uses the evidence already available. A controlled retry might recover from a temporary processing error. A different extraction configuration might help detect a clearly visible field. In either case, record the new processing run and candidate version so a reviewer is not unknowingly judging a result that changed underneath them.

Requesting evidence changes the submission. It needs a clear question, a destination and a way to associate the reply with the unresolved item. “Please upload the missing second page containing the declaration” is more useful than “Try again.” Avoid requesting an entire sensitive document collection when a narrowly scoped replacement would satisfy the approved requirement.

Do not present an inferred value as a repair for missing evidence. A model may produce a convincing amount, date or identifier from patterns in similar documents. That answer does not turn an absent source into a supplied source. If inference is permitted for a low-consequence draft, label it explicitly and prevent it from quietly satisfying a field that requires documentary support.

Put limits on repeat attempts. After the allowed processing attempts produce the same unresolved observation, route the item to clarification or an owner decision. Record why further processing is unlikely to change the evidence. This keeps queue effort focused on an action that can actually resolve the uncertainty.

Bind replacement evidence to a new review revision

Imagine a synthetic application with a missing declaration page. A reviewer requests the page, then another employee changes the applicant's name in the destination record. The replacement arrives with a different name. Even if extraction now succeeds, the earlier review context is no longer enough to accept the record update.

Create a new review revision that names the relevant submission set, destination record revision and applicable rules. The reviewer sees the difference and decides under current conditions. If the decision only authorizes preparing a draft, preserve that boundary when the workflow later reaches an execution service.

Handle duplicates explicitly. A replacement uploaded twice should not create two independently actionable approvals for the same intended effect. Keep submission identity separate from document content similarity: two identical files can represent a repeated upload, while two different files can belong to the same unresolved request.

When clarification expires or the submitter withdraws, transition the item out of the actionable queue without declaring it successfully resolved. Preserve the reason and any approved retention rules. An administrative closure and a supported acceptance are different outcomes, even if both reduce the visible queue count.

Test absence and contradiction as separate fixture families

Build independent fixtures for a missing page, an unreadable value, a contradictory pair of sources, a failed processing job and an ambiguous candidate. Define the expected route before running the extractor. A test that labels every unresolved item “incorrect extraction” cannot establish whether your workflow requests the right information.

Google Document AI's evaluation documentation explains extraction precision and recall. Those metrics are useful for extraction evaluation. The additional routing tests proposed here ask whether the application responded appropriately to the evidence available.

For every fixture, inspect the reason code, owner, reader-facing explanation and allowed next action. Confirm that no unsupported value satisfies a required-evidence rule. Confirm that a retry does not erase the prior failure and that a replacement source invalidates any review decision bound to the earlier submission revision.

Add one deliberately difficult fixture where reviewers cannot determine the authoritative source. The expected result may be a documented hold and escalation, not acceptance. A good test suite permits that outcome. Forcing every example to end in a populated record teaches the implementation to hide unresolved uncertainty.

Measure resolution, not just queue clearance

Separate clarification requested, awaiting source, resolved with evidence, administratively closed and still disputed. Track elapsed age as well as reviewer handling time. A missing attachment may take little reviewer effort but several days to arrive. Reporting only handling minutes hides that delay from the business owner.

In a hypothetical week, 40 items enter clarification. Twenty receive adequate sources, ten are withdrawn and ten remain unanswered. Calling all 30 departures “successful resolution” would combine supported decisions with withdrawals. Report the outcomes separately, then decide whether the wording of the request or the intake requirements needs improvement.

Inspect reopenings too. If a supposedly resolved item repeatedly returns because the wrong source was accepted, throughput has overstated useful work. Link the reopened item to the earlier decision rather than counting it as an unrelated new exception. That connection lets the team examine the procedure that allowed the premature closure.

Start with an exception record a reviewer can use

A useful record contains the item identifier, submission revision, required-evidence rule, observed gap, reason code, destination record revision, accountable owner, requested next action and decision history. Include only the source context necessary for the role. A queue needs enough information to resolve work without becoming an uncontrolled copy of every sensitive document.

Write one concise explanation per reason. “The submitted pages do not include the required declaration. Upload that page or ask the intake coordinator about an approved alternative” tells the reader what to do. Avoid stating that the AI “cannot understand” the document when the actual problem is an absent page. The right explanation reduces needless resubmissions and sets an honest expectation.

Use the review-queue pilot to test the record and pause conditions before connecting real effects. Read the field-confidence guide for the separate question of accepting an extracted candidate. For the wider design, see document-intake exception handling. You can also ask Ampity to review your workflow without turning access to the resource into a lead form.