Evaluate the Image Bytes Behind a Document Test
Bind document-image evaluation to the actual rendered input, preserve critical-field context and keep paired variants together before comparing extraction outcomes.
1. Compare the stimulus before comparing the answer
A text-only document test cannot establish how an image-input path handles the same document. Record the source revision, the locally rendered image and the exact bytes prepared for the request. Check whether cropping, scaling or page selection removed the evidence required for the task. Only then decide which image-quality and layout slices belong in an extraction comparison.
The reader here is an application or evaluation engineer moving from a clean-text prototype to page-image input. The bounded task is preparing an invoice draft's reference, amount due and currency, with no payment, supplier-master change or external write. The examples are fabricated teaching records. No model was called, no provider response was obtained and no accuracy, latency, savings or customer result is measured.
Existing guidance already covers field-level acceptance, missing evidence and slice-level release gates. This article resolves an earlier comparison question: did the image path receive the evidence that the text path and reference label assumed? A model comparison is confounded when the text includes currency that the submitted crop excludes, even if both cases have the same invoice ID in a spreadsheet.
There are three different judgments to retain. The stimulus record establishes which local file was prepared. A reference reviewer establishes what the visible evidence supports. The evaluation compares a supplied candidate result with those labels. None establishes that a business fact is true or that the resulting draft may cause a consequential action. Keep the task at draft preparation until its separate authority requirements are met.
2. Name one documented image route
The documentary candidate checked on October 9, 2026 is Claude Sonnet 4.5 on the Bedrock Runtime Converse API. Its base identifier is anthropic.claude-sonnet-4-5-20250929-v1:0; the proposed route is the US inference profile us.anthropic.claude-sonnet-4-5-20250929-v1:0 from us-east-1. The AWS model card lists image input and Converse support, and geo inference rather than direct in-Region support from that source. This is documented candidate scope, not observed account access or routing. Account prerequisites and processing-location admission remain separate reviews.
For Converse, the Message contract allows images only in user messages and lists up to 20 images, each no larger than 3.75 MB with height and width no larger than 8,000 pixels. Those are admission bounds, not a promise that tiny type will be recognized. This specimen prepares one PNG, avoiding mixed document blocks, animation and multipage assembly.
The ImageBlock reference names PNG, JPEG, GIF and WebP formats. The ImageSource reference separates bytes from S3 location as a union, and says AWS SDK callers do not need to base64-encode raw image bytes themselves. Here the source is local PNG bytes prepared for a hypothetical request. SVG is retained only as editable author source. No SVG is presented as a supported Converse image payload; no SDK request builder or service validator is implemented.
Do not transfer these restrictions to every Bedrock model or API. A provider's general vision page may discuss another API's image counts, file upload or transformations. Validate the selected route's contract before adopting those features. Document a contradiction instead of silently selecting the larger limit. Neither a catalogue entry nor a syntactically plausible payload proves account entitlement, permission, useful output or an unchanged future contract.
3. Keep source, local stimulus and provider processing separate
The source file may be a PDF, camera image or an authored synthetic page. A renderer chooses pages, dimensions, orientation and color treatment. A cropper or compressor can then create a different image. Record each transformation in order and hash the bytes after the last local transformation, not only the uploaded original. Otherwise two runs can cite one source hash while sending different visible evidence.
In this offline specimen, one editable 1,600 by 1,200 SVG page is locally rasterized three ways: a full PNG, a 160 by 120 PNG, and a 430 by 140 viewport that shows the numeric value while excluding the currency and its amount-due label. These files were rendered locally and inspected. They are synthetic stimuli, not scans from a customer. The one-tenth rendering changes character scale; the crop changes evidence. Neither represents an observed provider transformation.
Anthropic's current vision guidance cautions about unclear or small text, lost crop context and compression artifacts, and describes model-dependent resizing. That supports inspecting the actual locally prepared image. It does not reveal the pixels internally processed by this hypothetical Bedrock request. A local byte hash establishes reproducible local identity, not an immutable guarantee about provider preprocessing. Retain that uncertainty in the result rather than calling the local PNG the exact internal model stimulus.
Keep the original page available to qualified reference reviewers under the dataset's access policy. That does not mean secretly giving the candidate additional evidence. If the candidate gets only a crop, grade what that crop supports. A reviewer may know the original currency while the candidate legitimately cannot determine it. Record both source truth and stimulus-supported truth; using them interchangeably rewards guessing and penalizes correct abstention.
4. Work a context-loss counterexample
The fictional source prints reference INV-E27. Its subtotal is 1,000.00, tax 250.00, and the row labelled Amount due says EUR 1,250.00. The draft contract requires the reference plus the amount/currency tuple. Monetary normalization is explicitly two decimal places for this EUR-only teaching example, so the source amount due is 125,000 minor units. This is not a currency conversion or a general money parser.
The full local image preserves those distinctions. A fabricated candidate returning INV-E27, EUR and 100,000 minor units has copied the subtotal into the amount-due field. Two of its three required values match, but the accepted tuple is wrong. The error is consequential even though the reference and currency are correct. Adding many correctly copied descriptions would increase an all-field average without fixing the draft's amount. Report the wrong accepted tuple separately.
The isolated-number crop displays 1,250.00 but no invoice reference, currency or amount-due label. Under this contract the expected result is a deferred tuple. Copying the digits into an amount-due slot and inferring EUR from another test case is unsupported. Abstention is useful here because the submitted stimulus cannot establish the required relationship. The original page has evidence; the crop does not. That is a local preparation failure for a whole-invoice extraction comparison, not proof that the model cannot read currency.
The tiny rendering creates a different condition. Its labels have not been independently adjudicated for the selected task at native size. The specimen records those fields as unadjudicated. They do not enter a correctness denominator or become safe merely because an author can enlarge the image. An appropriately qualified reviewer must resolve that slice's reference state before a future candidate result can be scored. UNKNOWN is preferable to inventing a degraded-image gold value.

The unchanged full PNG is 1,600 × 1,200 pixels. This reader preview fits the article width; the companion contains the original bytes. No model received this file.

The unchanged crop is 430 × 140 pixels. Digits survive, but their required relationship does not. It is scaled down only when necessary to fit the reader; source bytes remain in the companion.

The unchanged tiny PNG stays at 160 × 120 CSS pixels, including in the expanded view. Its reference remains unresolved. Enlarging it would not independently adjudicate the original stimulus.
| Local stimulus | Reference condition | Supplied candidate | Meaning of outcome |
|---|---|---|---|
| Full page | Three required values supported | INV-E27, EUR, subtotal 100,000 | Two values match; wrong accepted amount tuple blocks the supplied comparison |
| Number-only crop | Required relationships unavailable | All three values deferred | Supported abstention; cannot certify whole-page extraction |
| Tiny page | Native-size reference unresolved | All three values deferred | Three unadjudicated fields; no correctness claim |
- Full page
- Three values supported. Supplied INV-E27 and EUR match, but 100,000 is the subtotal. Wrong accepted amount tuple blocks the supplied comparison.
- Number-only crop
- Reference, currency and amount-due relationship unavailable. All three deferred values are supported abstentions, not a whole-page extraction pass.
- Tiny page
- Native-size reference unresolved. Three deferred fields remain unadjudicated and outside the correctness denominator.
5. Pair variants without leaking their family
Assign a source-family identifier before generating variants. The full page, tiny page, number crop, OCR transcript and any prompt example containing this same invoice belong to E27's family. Keep that family in one development or held-out partition. Randomly splitting filenames can place the full source in development and its crop in the purported holdout. Different bytes do not make their underlying facts independent.
Use paired variants to diagnose preparation, not to manufacture sample size. Three renders of E27 are one source family with three stimulus conditions. They provide a useful counterexample but no estimate of performance on future invoice families. A later authorized dataset needs actual population coverage, document rights and independently administered labels. The review-queue pilot owns that broader operation.
Freeze labels before inspecting candidate answers. For each required field specify visible/supported, unavailable or unadjudicated, together with source evidence and permitted normalization. Require adjudication when reviewers disagree about whether a crop contains enough context. The model's plausible explanation cannot supply the reference label for itself. Nor should a failed crop be relabelled outside scope after results arrive without recording a new comparison version.
Compare paths on the same task with explicitly disclosed stimuli. A text transcript retaining all table labels is a different condition from OCR that loses row association. A crop-plus-full-page path is a different condition from a crop alone. These may be worthwhile alternatives, but a fairness record must name the extra evidence and processing effort. Keep timeout, invalid input and missing result as observed failure states in any actual future run, rather than dropping them from the tested population.
6. Retain an inspectable stimulus-equivalence record
The filled record below describes the offline E27 example. Replace it with permitted references for a real evaluation. It is a reusable artifact, not an AWS request schema or a release sign-off.
- Task and prohibited effect
- Prepare INV-E27 reference and EUR amount-due draft tuple. No payment, master-data change, message or external write.
- Source family and split
- E27 synthetic family, heldout-teaching partition. Full, small, crop and source text remain one family; no independent statistical holdout claimed.
- Source and local image identity
- Editable SVG and three PNG hashes/dimensions are recorded in the anonymous companion's manifest. Files were locally rendered, not transmitted.
- Transformation and lost context
- Full rasterization, one-tenth rendering, or number-only viewport. Crop removes reference, currency and amount-due label.
- Route and provider uncertainty
- Documentary Sonnet 4.5 Runtime Converse US-profile candidate. Account support and internal provider processing are unobserved.
- Reference and permitted normalization
- Full source supports INV-E27/EUR/125000 minor units. Crop's tuple is unavailable; tiny fields unadjudicated. Two decimal EUR-only normalization.
- Supplied output and slice disposition
- Full fabricated output substitutes subtotal 100000 and is wrong. Crop defers correctly; small remains unresolved. No measured model claim.
- Owner, stop and next evidence
- Evaluation owner holds comparison on wrong accepted tuple, identity drift, family leakage or missing reference adjudication; obtain independent review before any authorized trial.
The offline image-stimulus companion includes the actual synthetic PNGs, editable source, separate labels, fabricated outputs and small reusable evaluator. It checks local identity and grades arbitrary supplied outputs; it cannot see the request a service received or authenticate a returned answer. A fabricated result can score perfectly, which is why provenance must remain separate from arithmetic.
- Task and prohibited effect
- Name the accepted field tuple, downstream use, excluded actions and accountable owner.
- Source family and split
- Assign a family before rendering; bind all variants/transcripts to one partition and record any template-group leakage risk.
- Source and local image identity
- Record immutable revision references, source/render/final-local-byte hashes, dimensions, format, page order and actual request capture when authorized.
- Transformation and lost context
- Retain crop coordinates, scaling, compression and orientation versions; inspect whether required labels or relationships survive.
- Route and provider uncertainty
- Record exact model, endpoint/API/profile/source Region, checked contract, account prerequisites and what internal preprocessing remains unknown.
- Reference and permitted normalization
- Assign independent source and stimulus-supported labels, field states, evidence locations and exact permitted conversions; resolve disagreements.
- Supplied output and slice disposition
- Separate actual observed results from supplied fixtures; report wrong accepted critical tuples, abstention, unresolved labels and denominators.
- Owner, stop and next evidence
- Name a hold condition, preparation fix or clearer-source request, reviewer and separately authorized next action with version/recheck triggers.
7. Choose the alternative that repairs the evidence
If a crop removes currency, expand its context or use the full page; changing the prompt cannot restore omitted pixels. If the full page makes small text difficult, compare a context-preserving local rendering with a documented additional-region or page strategy. Record additional images and ordering, then recheck the selected request limits. Avoid treating more crops as free, independent evidence.
An OCR-first route may be preferable when stable extraction and spatial relationships are the primary requirement. Textract's input guidance discusses image quality and warns against unnecessary downsampling of supported files. That guidance belongs to Textract, not a universal Bedrock preprocessing rule. Its useful lesson for this comparison is to evaluate the chosen preparation path instead of assuming one universal low-resolution representation is fair to every candidate.
For specialized document extraction, Google Document AI's metric documentation compares predictions with annotated test documents and describes aggregate weighting by label occurrences. Do not copy that service's matching rules into this specimen by accident. Here correctness uses exact reference/currency/minor-unit values, with missing context and unresolved labels explicitly separated. That definition is intentionally narrower than document legitimacy or posting eligibility.
Cost may change with the selected route, modality and model, as the Bedrock pricing page states. No current per-image price, image-token count or saving is estimated here. A later paid comparison needs actual route-appropriate usage, repeats, preparation and review effort per accepted task. A smaller file is not proof of lower billed usage, and more correct easy fields are not proof of fewer expensive wrong drafts.
8. Hold conclusions outside the evidence boundary
This three-variant, single-family fixture demonstrates a confounded stimulus and a scoring distinction. It does not establish a model ranking, production slice frequency, rare-error rate, reviewer reliability, account eligibility or image capability on another route. Labels are author-supplied and need independent review. The tiny variant remains unadjudicated instead of being forced into a pass/fail answer. Inspecting these local stimuli does not establish that a service received or processed them.
File hashes detect accidental local changes only when the expected manifest is trusted. They do not authenticate a document's issuer, establish data rights or prove that a later service request used those bytes. A changed crop or source requires a new stimulus record and affected labels. A changed provider/model contract requires another documentary check; do not silently reuse an old result because the task name stayed the same.
Start with one existing text-only test and inspect the final local image at its real dimensions. Record which required evidence survived, place its variants in one family, and have a second reviewer challenge the label for the most consequential field. If that reviewer cannot establish the amount/currency relationship, hold the comparison and repair the source preparation before scheduling any separately approved evaluation. Ampity can help review that bounded evaluation contract; access to this guidance does not require contact details.
Related services
AI Observability, LLM Monitoring & Governance
LLMOps consulting for AI observability, LLM monitoring, evaluation and guardrails. Review production answer quality, operating failures and cost evidence.
AI Agent Development Services
AI agent development for enterprise workflows. Design scoped tools, permissions, evaluation and recovery before deploying autonomous or multi-agent systems.