AI Document Intake: Evidence, Exceptions and the Decision to Automate

Decide which document fields can proceed automatically and which need review. Design provenance, confidence evaluation, exception ownership, queue capacity and a...

audience="Application owners, operations leaders and engineers building document intake for invoices, onboarding packs, claims or service requests." decision="Which extracted fields may advance a business process, what evidence must accompany them, and how uncertain or conflicting records are resolved." position="Automate a defined decision, not an entire document merely because its average confidence is high. Preserve field evidence and make exceptions an owned workflow before permitting downstream writes." scope="A proposed design with an illustrative invoice example. It is not an Ampity customer result, a vendor accuracy benchmark or financial or legal advice." outputs={[ 'A field consequence and evidence register', 'An acceptance policy with an explicit review path', 'A version-bound adjudication record', 'A review-capacity calculation', 'A representative evaluation set', 'A controlled pilot and rollback checklist', ]} />

Executive summary

A document-intake system can read nearly every word correctly and still make the wrong business decision. An invoice total may be clear while the currency is missing. A bank account may be legible while belonging to a different supplier. A model may extract a delivery date accurately from an old revision. These are different failures, and a single document confidence score cannot tell an operator which one occurred.

The central recommendation is to make the acceptance unit match the business consequence. Preserve the original document revision, the evidence for each consequential field and the policy that allowed it to proceed. Route absent, unreadable, contradictory or unauthorized information into named exception classes. A reviewer should resolve the specific uncertainty, not simply click an approval button beside a model-generated summary.

This paper proposes an intake design that separates safe file admission, extraction, validation, adjudication and business execution. It uses an illustrative invoice workflow because the consequences are easy to distinguish, but the reasoning applies to other document-heavy operations. It does not claim that a particular cloud service achieves a particular accuracy, that human review eliminates errors, or that the proposed controls have been deployed for a customer. The useful outcome is a release decision supported by evidence: which narrow class of records can advance, under what conditions, with what remaining risk and who owns the exceptions.

Scope and operating assumptions

Assume documents arrive through an authenticated portal or a controlled integration. They can be digital PDFs, scans or images. A case can contain several documents and later revisions. The system needs to create or update a business record, but creating a candidate record does not necessarily authorize a payment, acceptance of a claim or activation of a supplier. Those later actions have separate authority requirements.

Assume the team can retain source revisions for an agreed period, restrict access by case and tenant, and assign an operational owner to unresolved work. These are prerequisites to the proposed approach. If reviewers cannot access the evidence, if the business cannot define required fields, or if no one owns the queue, adding a stronger model will not repair the process. Start by changing the operating arrangement rather than hiding the gap in a prompt.

The design does not assume all fields have confidence scores or all providers express scores on the same scale. It also does not assume documents are truthful. Extraction establishes what a supplied document appears to say. Verification against an approved supplier record, contract or independent authority establishes a different fact. Keep those statements separate throughout the implementation and in any customer-facing explanation of what the automation does.

Define the business decision before the schema

Begin with the smallest action worth automating. For an invoice, this might be creating a draft accounts-payable record with the document attached. It might be proposing a match against an existing purchase order. It should not automatically mean paying the invoice. These actions differ in reversibility, authority and exposure, even if they use identical extracted fields.

Write a short decision contract. Name the action, the recipient system, the required facts, the permitted source of each fact and the evidence of completion. Identify fields that must come from existing trusted records rather than the uploaded document. In the illustrative case, a document can supply an invoice number and billed amount; a supplier identifier may need resolution against an approved master record. A newly printed bank account should not silently replace the master record's payment instructions.

Then define the schema. Distinguish a candidate value from an accepted value and a missing value from a deliberately inapplicable one. Include a reason when a field is not required. This prevents a common shortcut: making every column optional to get imperfect extraction through the database. A permissive storage schema is sometimes useful for candidates, but it must not become the acceptance policy for business execution.

Classify field consequences individually

Build a field register around the proposed action. A descriptive memo can tolerate a different error than a currency or recipient. A low-risk field can still become high-risk when it participates in routing, identity or access. Treat consequence as a property of the field's use, not just its datatype. An ordinary string becomes consequential when it selects the supplier receiving money.

| Field in the illustrative invoice | Candidate source | Acceptance concern | Safe response to uncertainty | |---|---|---|---| | Invoice number | Visible source span | Missing, altered or duplicate identifier | Hold for identification or duplicate investigation | | Supplier identity | Document plus approved master data | Similar names or unapproved supplier | Resolve identity without changing master data | | Total and currency | Source spans and consistent line calculations | Decimal interpretation, currency omission, conflicting totals | Review the specific field and its dependencies | | Purchase-order reference | Source span plus order lookup | Wrong order or inaccessible order | Route to the owning business team | | Payment instructions | Separately controlled supplier process | New or changed destination | Exclude from this automation's write permission | | Descriptive memo | Source text | Meaning lost during summarization | Preserve source and label generated text |

Review the register when the action changes. A draft-import pilot cannot serve as approval for payment automation later. Record who accepted the consequence classification and what evidence would trigger reconsideration. The purpose is not to produce an impressive matrix; it is to stop fields from acquiring new authority simply because a downstream team finds them convenient.

Separate absence, unreadability and contradiction

At minimum, distinguish four evidence states: supported candidate, absent required evidence, unreadable evidence and conflicting evidence. Add explicit states for unsupported document types or unauthorized sources where needed. A blank value cannot describe all of them. If the currency is not printed, rerunning OCR may never solve the case. If two totals disagree, a higher score for one does not explain why the other is irrelevant.

These states need different review actions. Absence can require requesting another document. Unreadability can require a clearer scan. Contradiction can require a business owner to establish which revision or statement governs. An unauthorized source can require a separate verification process. Allowing reviewers to pick only “approve” or “reject” hides the work needed to make the record usable.

Preserve the unresolved condition when a reviewer cannot resolve it. Do not force a guessed value merely to empty the queue. A case that is waiting for a supplier response should remain distinguishable from one waiting for internal adjudication. Give each waiting state an owner and a next review time. Otherwise the intake dashboard can look healthy while consequential records quietly age beyond the deadline they were meant to support.

Preserve field provenance across revisions

The reviewer needs to find the source behind a candidate without reconstructing the entire extraction run. Keep a revision reference, a page reference, the relevant source region or text span, the extractor version, any normalization applied and the candidate's evidence state. If a provider does not supply spatial anchors, make that limitation visible. A generated explanation is not a substitute for the source itself.

For example, the Amazon Textract Block reference describes page, geometry and relationships associated with extraction results. It also says block identifiers are scoped to one operation. A proposed application record should therefore qualify provider identifiers with the extraction run and document revision rather than treat a block ID as a permanent business identifier.

Retain both the raw candidate and the accepted value. A reviewer changing “1,250” to a normalized numeric amount should record the interpretation and source, not overwrite the original extraction invisibly. If the source revision changes, do not repoint the old decision to the new file. Create a new candidate revision and identify which earlier decisions have become stale. Provenance is useful only when the evidence continues to refer to the version actually reviewed.

Use confidence as evidence, not permission

Provider confidence is one input to acceptance. It is not proof that a supplied invoice is legitimate, that a purchase order is approved, or that a user can initiate payment. The Textract best-practices guide recommends considering the use case's sensitivity when deciding how to use scores. That supports a consequence-aware threshold policy, not a universal percentage for every document.

The Document Intelligence confidence guidance also distinguishes levels of prediction confidence and notes that not every field returns a score. A confident cell can coexist with uncertainty elsewhere in its row. In the proposed design, an absent score is explicit missing evidence, not a value that is conveniently replaced with a high default.

Evaluate acceptance on representative labeled cases before choosing a cutoff. Separate field correctness from record completeness, identity resolution and business-rule validity. A record with an accurate amount but unknown currency must not inherit acceptance from its stronger fields. Keep score interpretation specific to the provider, model version, field and document class. When those change, reassess the policy rather than assuming the previous threshold still has the same practical meaning.

Evaluate the selective automation policy

Measure the quality of the records allowed through, not just the extractor's overall quality. The relevant population changes when uncertain cases are deferred. Report how many records were eligible, how many advanced automatically, how many were deferred and what errors remained among the accepted records. A policy can appear accurate by accepting almost nothing; coverage and delay expose that tradeoff.

The Google Document AI evaluation documentation explains precision, recall and confidence-threshold behavior using annotated test data. Its aggregate metrics are weighted by label occurrences, and its threshold optimization targets F1. A business acceptance policy can require a different objective. Frequent descriptive fields should not dominate the release decision for a rare but consequential payment-related field.

In a proposed evaluation, label critical fields separately and retain the raw document family, quality, language and revision characteristics needed to understand failures. Compare the selected policy against a review-first baseline on the same cases. Report uncertainty when the sample is too small to estimate a rare error credibly. Zero observed failures in a pilot is evidence about that pilot, not proof that the true failure rate is zero or that the next supplier's layout will behave identically.

Put deterministic checks around bounded extraction

Use AI where interpretation is actually required, such as variable layouts or ambiguous textual descriptions. Use explicit rules for checks whose meaning the business can define: required currency, allowed date ranges, arithmetic consistency, order existence and permitted action types. Keep the extractor from granting its own exemptions. A prompt asking it to “validate everything” cannot replace rules whose results must be reproducible.

Some integrity checks belong in storage as well as application code. The PostgreSQL constraint documentation explains that a check expression accepting null does not itself prevent null values and that cross-row restrictions require appropriate mechanisms rather than an ordinary row check. If PostgreSQL is selected for this proposed design, distinguish candidate storage from accepted-record constraints and use the mechanism that matches the invariant.

A successful arithmetic check is not proof of truth. Fraudulent totals can add up perfectly. A master-data match can be wrong if the matching rule chooses the nearest name without adequate identity evidence. Label each check with the fact it establishes and the facts it does not establish. This makes operational review more useful than a screen full of green ticks whose scope no one can explain.

Compare three operating approaches

Review-first intake sends every candidate through adjudication before a consequential write. It can be appropriate when document classes are unfamiliar, exposure is high or the review process itself is still being learned. Its cost includes staffing and delay, and it still needs quality control. Human review is not automatically reliable when the interface hides source context or the queue pressures reviewers to approve quickly.

Straight-through processing accepts a defined record class without routine review. It can be appropriate for well-understood sources, constrained actions and tested acceptance policies. It should still sample accepted records for independent checks and hold cases outside the release scope. “Straight through” describes a path, not a claim that exceptions no longer exist. The system needs a fallback when that path loses its assumptions.

Selective automation routes supported records through automatically and sends specific exceptions to review. This is the proposed starting objective after a review-first pilot, but not an unconditional recommendation. It adds routing, adjudication and versioning complexity. Choose it only when the accepted subset creates useful throughput and the remaining queue can be operated. If nearly every case needs review, improving document submission or source quality may be more valuable than maintaining a complex automation layer.

Work through one uncertain invoice

Consider a hypothetical case: a supplied invoice prints a total of 1,250, a supplier trading name and an order reference. The source does not identify a currency. The supplier master contains two legal entities using similar trading names. The order reference exists, but the document's line quantity differs from the order. These facts are invented solely to illustrate the decision, not taken from an Ampity engagement.

A naive path fills the currency from the operator's location, picks the nearest supplier name and imports the record because the printed total was confidently extracted. The proposed path records the total as a supported candidate, the currency as absent, the supplier as unresolved identity and the quantity as a business discrepancy. It can create a limited intake case, but it cannot manufacture an accepted invoice by filling the gaps with plausible defaults.

The reviewer opens the cited page and confirms that the currency is genuinely absent rather than missed by extraction. The supplier owner resolves legal identity using the approved process. The purchasing owner determines whether the quantity difference is acceptable or requires correction. Each resolution records its source and authority. Only then can the accepted record advance to the action permitted by this workflow. The payment boundary remains separate, including any required approval and destination controls.

Design an adjudication record, not an approval flag

Record the question being resolved, the evidence considered, the selected value, the reviewer, the decision time and the exact candidate revision. Include a reason code that distinguishes correcting extraction from accepting a business exception. Those decisions have different implications for model evaluation. A quantity discrepancy accepted under a purchasing policy must not be mislabeled as a model extraction correction.

Use a claim or lease mechanism to reduce conflicting reviewer work, but still reject stale submissions through revision checks. A lease expiring does not mean the first reviewer stopped thinking about the case. Another reviewer may have completed it. The submission should name the revision the reviewer saw and fail safely if the record changed. Show the updated evidence rather than quietly applying a decision to a different candidate.

Make correction and override visibly different. A correction asserts that the accepted value accurately represents the evidence. An override permits a defined action despite a known condition, under a policy and authority that allow it. Do not let a generic override button bypass identity, access or financial controls. Where an exception is not overridable, the interface should explain the required next step instead of presenting a misleading approval option.

Calculate review capacity before promising speed

Estimate incoming cases, deferred share, review time by class, specialist availability and the deadline distribution. An average handling time can conceal rare cases that require a different team. Track active review time separately from waiting for external evidence. The former consumes reviewer capacity; the latter still contributes to customer delay and the outstanding-case inventory.

For an illustrative planning calculation, suppose intake receives 1,200 cases per working day, 25 percent are deferred and active review averages six minutes. That creates 300 reviews and 1,800 minutes, or 30 reviewer-hours per day. If one reviewer has six productive review hours available, five reviewers merely match the assumed average daily workload. This is not a safe staffing recommendation: it includes no spare capacity for variation, rework, specialist absences or backlog recovery.

Changing the threshold can increase review arrivals quickly. Test the queue effect alongside extraction quality before shipping a stricter policy. Decide what happens when capacity is exceeded: reduce the eligible scope, suspend automatic writes, prioritize explicit deadlines or request better source documents. Do not solve overload by hiding deferred cases, automatically approving them after a timeout or calling unreviewed records “processed” in a dashboard.

Build a reviewer interface that exposes uncertainty

Show the source region beside the candidate and its reason for review. Let the reviewer expand to the full page and related revisions when context matters. Mark normalized values and generated summaries as transformations. A cropped amount without its currency context can encourage the wrong decision even when the OCR itself was correct. Preserve enough surrounding material for the reviewer to understand the field.

Group related exceptions, but avoid making one approval silently resolve unrelated questions. An identity specialist resolving a supplier match should not automatically approve a quantity discrepancy. Present unresolved dependencies before the final eligibility decision. Make keyboard navigation, zoom, readable contrast and mobile limitations explicit in testing. If consequential review is not supported on a small screen, direct the user to the supported environment rather than pretending the source is readable at any size.

Do not preselect the model's answer as if agreement is the default. Measure reviewer disagreement and correction patterns, and independently inspect a sample of approvals. A fast reviewer can appear productive while missing systematic errors. Explain the limits of the review role and show how to escalate uncertainty without penalty. The interface should help people make a defensible decision, not merely improve the rate at which they close queue items.

Admit uploads through a separate security boundary

Operational ownership must include the security boundary: a quarantined file needs a visible status, restricted access and an owner who can request a safe replacement. Keeping it out of extraction protects downstream processing, but silently dropping it would leave the intake obligation unresolved.

The uploaded file is untrusted before any model sees it. The OWASP file-upload guidance recommends layered controls, including allowed types, size limits, validation beyond a client-supplied content type and controlled storage. Treat file admission as a separate implementation responsibility, with restricted parsing and evidence of the controls actually applied.

After admission, document text remains untrusted input. An instruction embedded in an invoice is part of the document, not authority to change the workflow. In the proposed design, extraction has no credentials for consequential business writes. It returns constrained candidate fields and evidence references. Server-side validation and authorized execution remain separate. This separation limits the effect of a generated answer that misunderstands document content, but it does not remove the need to test model behavior and parser security.

Restrict source access by tenant and case. Check authorization when evidence is opened, not just when a case is initially created. Avoid unrestricted public URLs in review notes or support tickets. Decide retention, deletion and audit-access requirements with the responsible business and privacy owners. The technical design should enforce the resulting policy without claiming that one retention duration is legally appropriate for every document or jurisdiction.

Minimize data in models, logs and review exports

Inventory which document fields are necessary for the decision. A supplier-onboarding pack may contain personal identifiers that an invoice-matching task does not need. Avoid sending entire case histories to a model simply because retrieval makes them available. Limit access to the source needed for the current extraction and adjudication question, and verify the selected provider's current processing and retention arrangements before using sensitive production material.

Operational logs should establish continuity without becoming a second uncontrolled document archive. Record case IDs, revisions, routing reasons, versions and timing. Protect raw text and source regions separately. A hash can help bind a decision to a specific artifact, but it does not by itself provide confidentiality, establish authenticity or make a document irreversibly anonymous. Define access and retention for evidence and operational metadata independently.

Review exports and evaluation sets also need ownership. An engineer downloading failed examples creates a new copy outside the normal application controls. Use approved, minimized datasets and record where they can be stored. Avoid assuming that a test environment is harmless because it cannot trigger payments. It can still expose confidential source documents, personal data or commercially sensitive supplier information.

Bind acceptance to versions and current authority

An accepted record needs the source revision, extraction version, validation-policy version and adjudication revision that supported it. Before a consequential write, verify that those references still satisfy the current action contract. A model update does not need to invalidate every historical record, but the team must decide whether a new policy requires re-evaluating records still waiting for execution.

Check business authority at execution time. The actor may have lost access, an order may have closed, or the supplier may have been suspended after review. A previously correct extraction does not authorize action against a changed business state. Preserve the earlier decision as history while holding the new action for the appropriate check. This is an application responsibility; the extractor cannot establish current permissions from a document snapshot.

Create an operation identity for any external write and preserve it through retries. A lost acknowledgement can leave an accepted record in an uncertain execution state. The AI action-recovery paper examines that separate problem. Do not reopen extraction merely because the business write timed out. Re-extracting can generate a new candidate, conceal the original unresolved operation and make a duplicate action more likely.

Reprocessing and duplicates need different identities

Keep intake case identity, source-artifact identity, extraction-run identity and business-operation identity distinct. The same document may be extracted twice for a legitimate model comparison. The same invoice may also arrive as a scan and a digital PDF. Identical bytes are one useful duplicate signal, but different bytes do not prove different business obligations. Deduplication must reflect the action and source semantics.

A proposed duplicate investigation can consider supplier identity, invoice reference, currency and amount, with explicit handling for reused numbers or amended invoices. Do not automatically merge ambiguous cases based on approximate text matching alone. Preserve why the system considered them related and let the owning process resolve whether one supersedes another. A false merge can omit a legitimate obligation just as a missed duplicate can create a repeated one.

Reprocessing should produce a new candidate revision with a relationship to its predecessor. Compare changed critical fields and policy outcomes before allowing replacement. If an old revision has already caused a business effect, correction requires the downstream process's supported amendment or recovery path. Deleting the old extraction record is not a reversal. Preserve enough history to explain both what happened and why the new version differs.

Test failure paths and evaluation integrity

Construct tests around failures that cross component boundaries. Include a multi-page document with a missing page, a revised file uploaded while review is open, two reviewers submitting the same revision, a parser timeout, a provider outage, a conflicting supplier match and a business write with a lost response. Verify resulting states and owners, not merely the status code returned by each API.

Keep the evaluation set separate from examples used to adjust prompts, layouts and thresholds. Otherwise repeated tuning can make the visible score improve without improving behavior on new documents. Retain difficult cases instead of removing them as outliers unless they are formally outside scope. Report that scope so operations knows where new arrivals will go. A supplier's unfamiliar layout is not outside scope simply because it lowers a presentation metric.

Review critical-field failures individually. Ask whether the failure came from admission, extraction, normalization, source ambiguity, identity resolution, policy, reviewer interaction or execution. Assign remediation to the responsible layer. Replacing the model may help extraction, but it will not repair an approval attached to the wrong revision or a queue with no owner. The acceptance report should show both observed evidence and what remains untested.

Operate degraded modes deliberately

When extraction is unavailable, retain admitted cases and show that they await processing. Use bounded retries appropriate to the provider and preserve run identifiers. Avoid switching to a different model or region without the agreed data-handling and evaluation checks. A fallback can have different capabilities and processing terms, so it belongs in the design and release evidence rather than an emergency prompt edit.

When review is unavailable, decide which already-tested low-risk classes may continue and which must pause. When the destination system is unavailable, keep accepted candidates separate from unresolved writes. Monitor oldest case age by reason, not just queue length. A stable total can hide one critical case that has been stuck for days while easier cases move through.

Assign owners for provider incidents, review capacity and business exceptions. Give each owner a concrete runbook with a supported action, evidence to inspect and a condition for returning to normal operation. After recovery, reconcile intake, adjudication and execution counts. A cleared worker queue does not prove all source documents received a business outcome or that no action was repeated during recovery.

Release a narrow pilot before expanding scope

Start in observation or review-first mode on an agreed document class and action. Compare candidates with independently established outcomes. Build the field register, exception taxonomy and reviewer workflow while the scope is small. Then propose selective automation for a defined subset. Record the acceptance evidence, operating capacity, known limitations and the person authorizing the scope change.

The release checklist should include source-access controls, named exception owners, reviewer revision checks, duplicate behavior, uncertainty after a write and a demonstrated pause mechanism. Identify how new arrivals behave after automatic execution is disabled. A rollback should stop new eligibility decisions without erasing accepted records, ongoing reviews or unresolved operations. Confirm that operations can continue manually using the preserved evidence.

Expand one meaningful variable at a time: a new supplier group, document language, layout family or action consequence. Inspect its evaluation and queue effects before treating it as equivalent to the earlier scope. The goal is not to celebrate the number of documents touched by AI. It is to advance useful business work while keeping the evidence, authority and remaining uncertainty understandable to the people responsible for the result.

Review questions and limits

Before accepting the design, ask which facts come from the document and which require another authority. Ask what an absent field means, where a reviewer sees the original evidence, how stale decisions are rejected and who handles cases that cannot be resolved today. Ask which classes may proceed during an outage and how an uncertain business write is reconciled. These questions should have observable answers in the pilot, not only paragraphs in a proposal.

The proposed approach has costs: evidence storage, reviewer training, version management, evaluation maintenance and exception ownership. It cannot prove document truth, guarantee reviewer accuracy or remove every duplicate. Some workflows may be better served by improving structured submission instead of extracting unstructured documents. If a portal can require a verified supplier reference and explicit currency at source, that may remove uncertainty more effectively than another model pass.

Use this paper to prepare a decision packet, not to select a confidence cutoff in isolation. Include the action contract, field register, labeled evaluation slices, queue-capacity assumptions, review lifecycle and release evidence. Ampity can discuss a specific intake decision when you choose to contact us. Reading or downloading this resource does not require providing contact details, and the examples above are illustrative designs rather than customer case studies.