Pilot an AI Document-Intake Review Queue

Run a review-first document-intake pilot with field-level evidence, a reviewer worksheet, stale-revision tests, queue-capacity calculations and explicit conditions...

trigger="Document extraction produces plausible candidates, but reviewers cannot consistently explain why a record can proceed or where unresolved cases go." owner="The business process owner responsible for the accepted intake record." participants={['Application engineer', 'Review-team lead', 'Document specialist', 'Security reviewer', 'Evaluation owner']} prerequisites={[ 'One document class and one proposed downstream action', 'Approved synthetic or minimized pilot documents and independent reference labels', 'An isolated destination that cannot send messages, pay invoices or activate access', 'Named owners for extraction corrections and business exceptions', ]} outputs={[ 'A field consequence and evidence register', 'A completed adjudication worksheet and revision tests', 'A queue workload and delay report', 'A signed scope decision for a selective-automation pilot or a documented hold', ]} doneWhen={[ 'Reviewers can resolve each in-scope question from cited source evidence', 'Stale or conflicting reviews cannot authorize a new write', 'Unresolved cases retain an owner, reason and next review time', 'The eligible subset and remaining workload are measured separately', 'A pause rehearsal preserves the cases and decisions already in progress', ]} />

Choose a review-first pilot

Start with one document class and one action, such as creating a draft invoice record. Keep every consequential candidate in review while the team establishes its evidence and exception process. This playbook produces a scope decision for selective automation. It does not presume that automating every arrival is the correct result.

The procedures and numbers below are proposed examples, not Ampity customer results. Use a destination that cannot trigger real payments, notifications or access changes. If the team cannot isolate those effects, stop before replaying documents. A “test” invoice sent through a production integration can still reach a supplier, create an obligation or expose confidential information.

Record the pilot version and document population at the outset. A successful invoice pilot does not approve claim processing, supplier onboarding or a different language and layout family. The document-intake whitepaper explains the decision framework; the steps here define who gathers the evidence and how the pilot reaches an acceptance decision.

1. Write the action contract

Owner: business process owner. Output: a one-page action contract. Describe what the receiving system may do with an accepted record. Specify whether the pilot only creates a draft, updates an existing record or proposes a match. Name the actions it cannot perform. For an invoice draft, explicitly exclude changing payment instructions and initiating a payment.

List the facts required for that action and the authority that supplies them. The uploaded document can provide printed values; an approved supplier master may supply legal identity. The person responsible for purchasing may need to resolve an order discrepancy. Separate those responsibilities before choosing a schema or threshold, so the model does not acquire authority through an undocumented default.

Have the application engineer show the actual endpoint and its downstream effects to the process owner. Include webhooks and notifications. Acceptance requires agreement on the business effect, a named person who can pause it and a tested isolated destination. If the endpoint does more than the written contract permits, reduce its scope or choose another pilot action.

2. Assemble a representative document pack

Owner: evaluation owner, with the document specialist. Output: a versioned fixture manifest. Include normal records and deliberately difficult cases: poor scans, a missing required field, conflicting totals, unfamiliar layouts, a revised document, split tables and ambiguous supplier identity. Use synthetic cases where needed, and mark them as such. They test behavior; they do not establish the distribution of production arrivals.

Keep a separate representative sample from approved historical material when authority and handling arrangements permit it. Record document family, language, revision, relevant quality characteristics and why the case belongs in the pilot. Do not select only the examples that the current extractor already handles. Preserve rejected and unresolved records in the manifest instead of making them disappear from the denominator.

Assign immutable fixture identifiers. Retain reference labels separately from model outputs and record the expert who established them. Resolve disagreements about the expected outcome before using a case to grade extraction. Some cases may have a legitimate “cannot determine” outcome. Guessing a gold value would reward the system for guessing in the same way.

3. Verify admission and evidence access

Owner: security reviewer and application engineer. Output: an admission/access test record. Test the permitted file classes and the actual size, parsing and storage controls. The OWASP upload guidance describes layered checks and cautions against trusting the client-supplied content type. Keep admission failures outside ordinary extraction retries.

Test that a reviewer can open only the cases they are allowed to review. Change the reviewer's access after a case is queued and confirm that opening its evidence rechecks permission. Use controlled references rather than publicly shareable document URLs. Check that routine logs and error screens do not reveal source contents to a different tenant or unauthorized support user.

Make the quarantine response visible. Record a reason, an owner and a supported route for requesting a safe replacement. The operator should not need to copy a quarantined file into an unrestricted folder to keep work moving. Stop the pilot if the only available recovery bypasses the controls being tested.

4. Define field evidence and exception classes

Owner: document specialist and process owner. Output: a field register and exception taxonomy. For each required field, identify its consequence, source and dependencies. Distinguish supported, absent, unreadable, contradictory and unauthorized evidence. Include the value's original representation, normalization and source anchor where the selected extractor provides one.

Use this example register as a starting artifact, then replace the rows with the pilot's actual fields:

| Field | Required for this action? | Evidence needed | Exception owner | |---|---|---|---| | Invoice reference | Yes | Printed source span and duplicate investigation | Intake owner | | Supplier identity | Yes | Document evidence plus approved identity resolution | Supplier-data owner | | Total and currency | Yes | Separate source evidence for both and relevant arithmetic checks | Finance reviewer | | Purchase-order match | Contract dependent | Authoritative order lookup and discrepancy policy | Purchasing owner | | Payment destination | Excluded from this pilot | Separate controlled supplier-change process | Supplier-data owner |

Assign a resolution action to every exception class. Missing information may require another document, while a business discrepancy may require an authorized decision. Treat unreadability as an evidence problem rather than permission to infer a plausible value. The acceptance gate fails if reviewers must use an undocumented workaround to complete a required field.

5. Freeze the candidate and policy versions

Owner: application engineer. Output: a candidate-record contract. Bind each extracted candidate to the source revision, extraction run and applicable validation-policy version. Preserve raw values alongside normalized ones. Store whether confidence was supplied, rather than inventing a score for a field with no score.

The Document Intelligence guidance notes that score availability depends on the prediction type. The pilot must inspect the selected provider's actual response. The Textract guidance connects threshold selection to use-case sensitivity. Neither reference supplies a universal acceptance policy for this workflow.

Run the same fixture twice and verify that the two extraction runs remain distinguishable while referring to the same source revision. Then upload a revised file and confirm that the system creates new candidates. A previous decision must remain historical evidence about its original revision, not automatically become authority over the new one.

6. Give reviewers a usable worksheet

Owner: review-team lead. Output: reviewer instructions and completed sample adjudications. Show the candidate beside its source region with a way to inspect the full page. Highlight the reason for review and the fields that depend on it. A reviewer resolving supplier identity should not silently approve an unrelated quantity discrepancy.

Use the following record for each question. This is a worksheet contract, not a claim that these fields exist in every vendor's product.

Case / source revision:
Candidate revision / policy version:
Question and reason code:
Evidence opened and source reference:
Original candidate / accepted value:
Decision: correction | permitted override | unresolved | rejected
Authority for any override:
Reviewer / decision time:
Remaining questions and their owners:
Next review time if unresolved:

Train reviewers on the difference between correcting extraction and accepting a known business exception. Review an example where the source itself lacks currency and another where the extractor missed visible currency. These require different evidence and different remediation. Acceptance requires reviewers to explain the distinction using the source, without being coached to approve the model's candidate.

7. Keep the review and execution boundaries separate

Owner: application engineer. Output: an observed lifecycle trace. Trace one case from admitted evidence through candidate creation, review and eligibility. Record each version and owner change. The proposed view below focuses on the pilot's control boundary and the evidence returned to evaluation. It intentionally omits provider deployment details because they do not determine who may resolve a question.

Compare the trace with the written action contract. Verify that unresolved fields keep the case in an owned waiting state and that accepted extraction does not bypass the separate write permission. If a source revision changes, trace which eligibility decisions are invalidated. Do not reduce all these outcomes to a single “processed” flag.

8. Inject stale and concurrent reviewer submissions

Owner: application engineer, observed by the review-team lead. Output: reproducible state-transition tests. Open a candidate in two reviewer sessions. Let the first session complete a correction, then submit the second session's older decision. The older submission must be rejected or require explicit re-evaluation against current evidence. A warning that can be ignored without a fresh decision is insufficient.

Repeat with a new source revision arriving while review is open. Test a review claim expiring and another reviewer claiming the case. Lease expiry reduces scheduling conflicts; it does not prove the first reviewer stopped working. Compare the submitted candidate revision with the current one before committing the decision.

Record the final state, preserved decisions and explanation shown to the stale reviewer. Acceptance requires one understandable current decision and a history of the attempted stale change. Stop the pilot if the last network response wins regardless of evidence version, or if an old browser tab can authorize a write for a newly uploaded document.

9. Test the unresolved and unavailable paths

Owner: operations owner. Output: failure evidence and a recovery runbook. Use the missing-field fixture to put a case into waiting-for-evidence. Verify its owner and next review time. Supply a correction later and confirm that it resumes under the appropriate revision. Test an unresolvable case and record its supported disposition rather than forcing a guessed field to make the queue empty.

Simulate extraction being unavailable, review capacity being unavailable and the isolated destination rejecting a write. These are separate conditions. Admitted documents should remain discoverable when extraction is delayed. Accepted records should remain distinguishable from unresolved writes when the destination fails. The write-recovery playbook supplies a separate test procedure for uncertain external effects.

For each condition, name who investigates, which evidence they read and how they resume work safely. Capture the oldest affected case and the operator's ability to find it. A runbook passes only after an operator other than its author follows it. Do not infer successful recovery from the fact that a worker queue eventually empties.

10. Grade the accepted subset independently

Owner: evaluation owner. Output: field-level results and an acceptance-policy comparison. Compare extracted and adjudicated values with the independent reference labels. Separate extraction corrections, legitimate business overrides and unresolved cases. Review decisions may themselves be wrong, so sample them independently. Do not turn every reviewer approval into a gold label by default.

The Document AI evaluation documentation describes label-level precision and recall and how thresholds affect accepted predictions. Use those definitions where they fit, but also grade required-record completeness and business eligibility. A high aggregate score can coexist with errors in a rare critical field.

Report the candidate acceptance policy's automatic coverage, deferred share, accepted-record errors and delay. Keep a separate held-out set for the final policy comparison. Any adjustment made after inspecting its failures needs another appropriately independent check. Record sample size and scope; a small pilot with no observed error cannot certify a rare error rate or a document class it never saw.

11. Measure active work and waiting time separately

Owner: review-team lead. Output: a queue-capacity worksheet. Measure active review minutes by exception class, rework and specialist availability. Also measure elapsed case age, including waits for outside evidence. A supplier response can consume little reviewer time while still delaying an intake obligation for days. Both measurements belong in the operating report.

For a hypothetical workload of 600 arrivals per day, a 30 percent deferral rate produces 180 reviews. If measured active work were eight minutes per review, the workload would be 1,440 minutes, or 24 reviewer-hours per day. Four reviewers with six productive hours each would only match that average workload. The example includes no reserve for variation, absence, rework or backlog reduction and is not a staffing recommendation.

Record how a stricter acceptance policy changes deferral arrivals. Test at least one plausible high-arrival period using authorized fixtures. Set a capacity response before the pilot: restrict the eligible scope, pause new automation or reprioritize named deadlines. Do not let a backlog timeout turn into automatic acceptance. Acceptance requires an owned plan for the queue the policy creates, not just attractive extraction metrics.

12. Rehearse pause and rollback

Owner: application owner. Output: a tested pause record. Disable new automatic eligibility decisions while cases are in extraction, active review, waiting and accepted-but-not-executed states. Verify how each state behaves. Existing review work should remain visible, source evidence should remain accessible to authorized users and unresolved operations should retain their identities.

Restore the prior policy in the isolated environment and compare the resulting cases. Determine whether the prior version understands the new candidate schema and reason codes. If it does not, traffic rollback alone is insufficient; document the supported repair or migration path. Keep the version references needed to explain decisions made before the change.

Reconcile the fixture manifest with intake cases, decisions and isolated effects after recovery. Every supplied fixture needs an explained disposition. Stop if the team can only make the counts match by deleting unresolved records, merging uncertain duplicates or relabeling errors as exclusions after the fact.

13. Decide whether selective automation is ready

Owner: process owner, with engineering and review leads. Output: a scope decision and acceptance criteria. Select the document and action subset supported by the evidence. Describe it positively, including source types, required fields, policy version and remaining checks. Record what happens outside that subset. A decision to remain review-first is a valid pilot outcome when uncertainty or capacity remains unresolved.

Use this acceptance checklist before enabling a limited automatic path:

| Acceptance question | Required evidence | Owner | |---|---|---| | Can a reviewer establish each required field? | Completed source-linked worksheets, including ambiguous cases | Review lead | | Are stale decisions rejected? | Concurrent review and source-revision traces | Application engineer | | Are business authority and extraction separate? | Endpoint effect inventory and denied-action test | Process owner | | Can the queue be operated? | Measured work, age by reason and capacity response | Operations owner | | Is the selected subset defensible? | Independent evaluation with scope and sample size | Evaluation owner | | Can automation pause without losing work? | Reconciled pause/recovery rehearsal | Application owner |

Record failures beside passes. Do not approve a broader class because the pilot seems promising. Assign a review date or trigger based on changed inputs, model versions, policy, staffing or observed errors. The sign-off names what may proceed now and what requires another decision.

Finish with the pilot evidence packet

Put the action contract, fixture manifest, field register, reviewer worksheets, fault traces, capacity report and scope decision into a versioned packet. Link each accepted criterion to the record that proves it. Keep unresolved criteria and their owners visible. This lets another operator inspect the result without replaying a slide presentation or trusting the pilot team's memory.

Before the next scope expansion, choose one failed or deferred case and walk it through the process with the owning team. Improve the missing evidence or control, then repeat the affected tests. If you want Ampity to review that specific intake boundary, share the action contract and one minimized example through our contact page. Reading or downloading this playbook requires no contact details.