Why Document Confidence Scores Need Field-Level Review
Decide which extracted fields can proceed, which need review and how to test a confidence policy against errors, coverage and reviewer capacity.
A field score does not authorize a business action
Use document confidence scores to help route extracted candidates, not to authorize their consequences. A high score for a payment account is not evidence that the account belongs to the intended supplier. A low score for a delivery instruction does not necessarily invalidate an otherwise usable invoice. Review policy needs the field, its source evidence and the action that would consume it.
This article proposes a field-level acceptance method for an engineering team. The examples are hypothetical. They are not measured provider benchmarks, Ampity customer results or recommended payment controls. The business owner must define which decisions require independent verification and which can remain drafts.
The distinction matters because a document can be easy to read and still be wrong for the workflow. It may be a genuine invoice from an unrelated supplier, an outdated statement or a duplicate submission. Confidence in extraction addresses a narrower question than whether the document should cause a record change. Keep that wider question visible when discussing automation coverage.
Separate extraction, normalization and authorization
Treat a field as a candidate with provenance. Store the source document revision, page and evidence location alongside the raw text. Record the normalized value separately. If an extracted date is rewritten into the system's preferred format, the reviewer should be able to inspect both representations rather than seeing only a polished result.
Amazon Textract's best-practice guidance makes confidence thresholds dependent on use-case sensitivity. That is a reason to evaluate your field policy, not to copy a universal cutoff from another workflow.
Consider three decisions that often get collapsed into one green check. First, did the extractor produce a plausible candidate? Second, did normalization preserve its meaning? Third, is that value acceptable for the intended record and effect? A recognizer might confidently read “03/04/26,” while the application still needs a locale or another source to determine the date.
Give each decision a separate reason code. A reviewer rejecting an ambiguous date should not label it as an OCR character error. Otherwise the team may spend time replacing extraction software when the actual missing requirement is document-origin metadata. Improvements become easier to evaluate when failures are attributed to the step that introduced them.
Write policy around consequences, not one document average
Suppose a synthetic intake fixture contains an invoice number, total, currency, supplier identifier and bank account. A high average score can conceal a consequential field with weak evidence. Conversely, a harmless descriptive field can lower the average without affecting a draft's usefulness. The average does not tell you which consequence is acceptable.
For each field, write the allowed next step. “Populate a draft description” is different from “update supplier master data.” Where a field influences several actions, either evaluate each action separately or adopt the stricter policy explicitly. Do not let a low-risk first use make the value permanently trusted for a later high-risk use.
| Candidate field | Consequence in this hypothetical workflow | Evidence needed beyond a score | Proposed disposition | | --- | --- | --- | --- | | Invoice description | Editable internal draft | Visible source text and document revision | Allow draft creation with provenance | | Invoice number | Duplicate detection | Supplier identity and comparison with existing records | Hold when the identity or duplicate check is unresolved | | Total and currency | Posting proposal | Line arithmetic, currency evidence and record match | Review discrepancies before a posting proposal | | Bank account | Supplier master change | Separately approved verification procedure | Never approve the change from extraction confidence alone | | Delivery date | Scheduling proposal | Unambiguous date interpretation and relevant order | Ask for clarification if interpretation remains uncertain |
The table is a policy-writing example, not a production control specification. Its purpose is to show that the same extracted document can produce several different dispositions. Any execution service still needs current permissions, revision checks and the business rules applicable at dispatch.
Know what the provider actually scored
An application should not invent a confidence value when a provider does not return one. Microsoft's Document Intelligence confidence documentation describes score availability at different levels and notes that not all fields have a confidence score.
Preserve the provider response type and model version in your evaluation records. Distinguish a field score from a word score, a classification result and an application-generated heuristic. An internal label such as “trusted” can hide these differences from reviewers and downstream code. Prefer explicit labels that say what was observed and what was decided.
Do not assume that a value such as 0.97 means a 97 percent probability of business correctness. If that interpretation is important to a policy, test the relationship against independently labeled examples representative of the intended workload. Even a well-calibrated extraction signal cannot establish authority to change a payment destination.
Missing scores need an explicit path. Depending on the effect, that path might be review, source clarification or draft-only use. Assigning an optimistic default makes undocumented provider behavior part of your authorization logic. Assigning a pessimistic default without measuring the resulting queue can make the workflow unusable. Both decisions deserve visible ownership.
Evaluate errors and coverage together
Build a test set that includes clear documents, poor scans, changed layouts, ambiguous dates, repeated identifiers and missing fields. Keep its labels independent of the candidate being tested. Include examples from known difficult groups rather than allowing one frequent, easy layout to dominate the reported result.
For each candidate policy, record how many field values were accepted without review and how many of those accepted values were wrong under the labeling rules. Also record rejected correct values, unresolved cases and reviewer time. A policy that reduces accepted errors by routing almost everything to review has not demonstrated a workable automation benefit.
For illustration, imagine 200 labeled total fields. Policy A accepts 150, including three incorrect values. Its observed error fraction among accepted fields is 3 divided by 150, or 2 percent. Policy B accepts 100, including one incorrect value, giving 1 percent. B has lower observed accepted error but only 50 percent acceptance coverage instead of A's 75 percent. Neither result chooses the production policy by itself.
These small hypothetical counts are not a forecast or a statistical guarantee. Review the consequences of each error, uncertainty from limited samples and the additional workload. State whether “correct” means exact source transcription, correct normalization or valid business interpretation. Changing that definition between experiments makes the comparison misleading.
Inspect the difficult groups before choosing a threshold
Break down the evaluation by field, document family and relevant source conditions. A policy can appear strong overall while regularly mishandling a minority layout. An intake team needs to know whether that group should be excluded, reviewed or supported through a different extraction path.
Check paired fields as well. Correct totals with incorrectly inferred currency can still produce an unsafe posting proposal. An address split across several lines can be individually plausible but collectively belong to the wrong party. Evaluate the unit of information the downstream action needs, not only isolated cells.
Review selected errors with the people who understand the records. Ask whether a wrong candidate had obvious contradictory evidence, whether the interface hid a qualifier or whether the labeling rule was incomplete. Those questions point to different changes: a validation rule, better evidence presentation or a clarified field definition.
Retest after model, prompt, normalization or document-layout changes. Preserve the prior result and versions so a team can identify what changed. Do not silently relax a threshold because a queue has grown. Capacity pressure is useful evidence about the policy, but it is not evidence that a consequential value became safer.
Use a compact field-policy worksheet
Start with one field that affects a real decision. Record its definition, source formats, permitted normalization, consumer action, required independent checks, candidate acceptance rule, review owner and pause condition. Add the evaluation sample, accepted count, accepted errors and reviewer minutes. This gives engineering and operations one small artifact they can inspect together.
Include a worked rejection. For example: “Total candidate accepted by the extractor, but currency absent from the approved evidence. Posting proposal held; reviewer requests a complete source.” A concrete rejection proves that the policy is more than a confidence slider. It also lets the team test whether the interface explains what is missing without accusing the submitter of wrongdoing.
Record what happens when the reviewer cannot decide. A queue item needs a destination, such as clarification requested or specialist review, with an accountable owner. Repeatedly returning the same unclear document to the general queue creates work without changing the available evidence.
Use the document-intake review-queue pilot to rehearse this worksheet with synthetic records. The document-intake exception whitepaper covers the broader lifecycle, capacity and revision boundaries. If you want help examining a production workflow, tell Ampity about its constraints. Reading and downloading the guidance do not require an enquiry.