Compare Prompt-Cache Costs Without Hiding the Cold Requests

Compare direct Claude and Bedrock cache usage on the same ordered workload, including first writes, changed prefixes, idle gaps and rejected outputs.

Compare prompt-cache costs by replaying the same task arrivals and prefix revisions, then pricing the usage each route actually reports. Keep the first writes, quiet periods, rejected answers and retries in the record. An all-warm loop can explain a best-case path; it cannot establish the cost of the workload that will use it.

This article is for an application engineer comparing direct Claude with Amazon Bedrock. It provides a proposed measurement record and a small offline accounting exercise. The counters, task outcomes and cost units below are supplied synthetic inputs. They are not captured API traffic, current prices, an Ampity engagement or evidence that either route wins. No provider request was made.

1. Fix the API pair and the question before timing it

The documentary example uses direct Claude Messages with claude-haiku-4-5-20251001, versus Bedrock Runtime Converse with us.anthropic.claude-haiku-4-5-20251001-v1:0 from us-east-1. Both arms propose non-streaming text requests, one five-minute explicit checkpoint, no tools, no thinking and no batch processing. The source Messages reference identifies the dated model ID and block-level cache control. Do not substitute a different model alias while retaining the old comparison label.

AWS's Haiku 4.5 card documents Runtime Converse, this US profile and explicit caching with a 4,096-token minimum prefix. The card allows five-minute and one-hour TTLs. The selected source Region does not restrict inference to that Region. Those are documented capabilities, not evidence of account access, permitted input handling or current client behavior. Recheck the exact route before an authorized trial.

Decide which question the experiment answers. Holding a model route fixed while changing checkpoint placement tests that implementation change. Comparing two endpoints with their supported caching behavior tests route economics. Changing the model, prompt, task mix and TTL together produces a bundle comparison, not evidence that caching caused the difference. Name the changed factors so the reader of the result knows what can be attributed.

Use the existing accepted-task measurement method for the business denominator. This article isolates model input/output accounting. Review labor, retrieval, storage, network, integration and shared operations remain outside the numeric example. They still belong in a full adoption decision.

2. Treat a reusable prefix as an opportunity, not a hit

Provider prompt caching reuses eligible input processing. It is different from returning a previously generated answer. The semantic-answer-cache article owns freshness and permission checks for that separate mechanism. Here each task still needs its own accepted result.

In the direct Claude configuration, the marked prefix must match, including the content before the checkpoint. Anthropic's prompt-caching guide describes exact matching, five-minute refresh and the model's minimum length. It also says a cache entry becomes available once the first response begins. Simultaneously launching many copies of the first request is therefore a different experiment from sequential reuse.

The AWS cache guide explicitly warns that eligibility does not guarantee a hit. It also notes that cross-Region inference can increase cache writes under high demand. Record the counter change; do not diagnose routing, expiry or a provider fault from one miss alone.

For your experiment, bind a prefix revision to the actual permitted request-construction record. Put stable instructions and the approved reference material before the checkpoint; keep task-specific text afterward where that preserves the task. A revision label without the corresponding request evidence is weak: a framework could insert a timestamp before the checkpoint while the label stays unchanged. A digest is a comparison aid, not proof of authorization or of the provider's internal cache key.

Do not append irrelevant content just to cross a token minimum. That changes input, cost and potentially output behavior. If the useful prefix is too short, record that the proposed reuse is unavailable for this configuration. Consider another supported configuration only through its own quality and scope review.

3. Keep the order, gaps and warm-up costs

Freeze task IDs, arrival offsets, prefix revisions, expected input class, acceptance rule and observation cutoff before collecting results. Preserve tenant/workspace boundaries and planned concurrency. The same six questions delivered in a tight loop and across a quiet working day are not the same cache workload.

Keep two separate views when useful. A controlled sequential probe can establish whether the implemented cache controls ever produce recorded reads. A representative replay asks how often that happens under the intended arrivals. Report both by name. Do not copy the probe's read fraction into a capacity plan as though the replay measured it.

For a cold-start question, retain evidence about prior exposure rather than asserting that a new process has an empty provider cache. Anthropic documents no manual cache-clear operation. Waiting or using a new synthetic prefix can be part of a test design, but neither is a substitute for inspecting the returned usage. The first recorded request may already report reads; preserve that result and label its prior cache state unknown.

If you intentionally prewarm, treat it as an attempt with cost, timing and a stated allocation to the measured population. It does not complete a user's task. Include it in the relevant numerator and retain a separate warm-up identifier. Likewise, do not erase a failed attempt when a later retry succeeds. Compare first-attempt behavior and the complete operating record separately if the question requires both.

4. Normalize usage without counting the same tokens twice

For the selected Claude Messages shape, retain ordinary input_tokens, cache_creation_input_tokens, cache_read_input_tokens and output_tokens. The first field excludes cache reads and writes. The cache guide gives total input as the sum of the three input categories. Do not charge that sum at the ordinary rate and then add the cache charges again.

For Converse, keep inputTokens, cacheWriteInputTokens, cacheReadInputTokens and outputTokens. AWS's cache-specific accounting guidance likewise separates the three input categories. The TokenUsage reference makes cache fields optional. A missing field in a log projection is not evidence of zero consumption. Preserve the original permitted response and record how the adapter established an explicit zero or an unknown.

AWS also defines cache-write detail by TTL. If a request mixes durations, an aggregate write count is insufficient to assign different write rates. The companion deliberately supports one five-minute category only; it holds mixed or unknown TTL evidence. Extending the model requires new disjoint categories and reconciliation, not charging both an aggregate and its breakdown.

There is a documentation-reading trap: the generic TokenUsage description calls inputTokens tokens sent to the model, while the cache guide explains its narrower meaning with caching enabled. Use the cache-specific equation for this scoped interpretation and confirm it against real usage and billing before making a financial claim. Do not infer the meaning of another API's similarly named field.

The normalized arithmetic uses four quantities: ordinary input U, five-minute writes W, cache reads R and output O. At rates per million tokens, the model-cost estimate is (U × ordinary rate + W × write rate + R × read rate + O × output rate) / 1,000,000. A write rate is the full rate for that category, not necessarily a surcharge to add to an already charged input token. Keep actual currency, effective date, platform, route and contract basis with any real rates from Bedrock pricing and the source provider. The following cost units are not copied from those rate cards.

5. Work six arrivals, including the inconvenient ones

Suppose six distinct tasks ask for source-linked answers from a fabricated operations manual. The task contract requires the applicable manual revision, correct cited facts, no unsupported additions and completion within the owner's deadline. Five tasks are stipulated accepted in each arm; task Q5 is rejected in both. Its usage still counts. The exercise does not contain model outputs or establish those judgments independently.

Each request is assigned an 8,192-token reusable prefix, 256 ordinary input tokens and 128 output tokens. Those equal sizes are synthetic conveniences, not a claim that the two endpoints tokenize or answer identically. A real comparison retains each route's actual counts for the same semantic inputs. P1 and P2 are distinct manual revisions, not public policy or real customer documents.

Task and offsetPrefixDirect arm: supplied prefix usageBedrock arm: supplied prefix usage
Q1 at 0 minutesP18,192 write; 0 read8,192 write; 0 read
Q2 at 2 minutesP10 write; 8,192 read0 write; 8,192 read
Q3 at 4 minutesP10 write; 8,192 read8,192 write; 0 read
Q4 at 5 minutesP28,192 write; 0 read8,192 write; 0 read
Q5 at 7 minutesP20 write; 8,192 read0 write; 8,192 read
Q6 at 18 minutesP18,192 write; 0 read8,192 write; 0 read
Q1 at 0 minutes, P1
Direct: 8,192 write; 0 read. Bedrock: 8,192 write; 0 read.
Q2 at 2 minutes, P1
Direct: 0 write; 8,192 read. Bedrock: 0 write; 8,192 read.
Q3 at 4 minutes, P1
Direct: 0 write; 8,192 read. Bedrock: 8,192 write; 0 read.
Q4 at 5 minutes, P2
Direct: 8,192 write; 0 read. Bedrock: 8,192 write; 0 read.
Q5 at 7 minutes, P2
Direct: 0 write; 8,192 read. Bedrock: 0 write; 8,192 read.
Q6 at 18 minutes, P1
Direct: 8,192 write; 0 read. Bedrock: 8,192 write; 0 read.

Q3 deliberately defeats the assumption that an eligible repeat must hit. Its cause is unspecified. Q4 changes the prefix. Q6 revisits P1 after a long gap. These conditions make writes worth investigating, but the ledger uses the supplied counters, not a timer-based simulation that declares what AWS must do. Do not call these six rows a measured hit-rate comparison.

Use fictional rates of 2 ordinary-input, 3 write, 0.25 read and 8 output cost units per million tokens in both arms. Direct totals are U=1,536, W=24,576, R=24,576 and O=768. The four costs are 0.003072, 0.073728, 0.006144 and 0.006144 units, totaling 0.089088. Bedrock has W=32,768 and R=16,384, with the same U and O. Its total is 0.111616 units. The 0.022528 difference is entirely the substituted Q3 category under these supplied rates.

The corresponding model-only costs per stipulated accepted task are 0.0178176 and 0.0223232 units, each divided by five rather than six. The arithmetic shows how a changed cache mix affects this synthetic record. It does not establish lower direct-provider pricing, worse Bedrock performance or a production recommendation.

Now replace every prefix use with a read while leaving U and O unchanged. The resulting 0.021504 units describes an all-warm counterfactual, not either ledger. Dropping the initial and replacement writes would make that attractive number appear observed. A separate all-ordinary-input counterfactual is 0.107520 units. Neither counterfactual proves what an uncached invocation would produce, nor that omitting explicit controls disables every caching mechanism.

6. Make the offline checks fail for bad evidence

The offline cache-accounting companion (ZIP) links the standalone Python checker, synthetic packet and tests. It uses exact rational arithmetic. It checks the supplied task and attempt manifests, required native usage fields, one TTL, category totals, complete task dispositions and the common comparison rule. It reports SUPPLIED_LEDGER_ONLY, never production approval.

Remove Q1 from an arm while retaining its attempt manifest and the check must hold. Replace a missing read count with an absent field and it must hold, not invent zero. Duplicate an attempt or task disposition, join an attempt to a different prefix revision, change one arm's acceptance rule, or supply negative/boolean token counts: each should stop the calculation. A nonzero cost with zero accepted tasks has no unit-cost ratio; the cost is retained.

The checker cannot know whether a supplied traffic manifest is exhaustive. Deleting both an attempt and its declaration can hide real traffic from this offline model. That is why the real evidence plan needs an independent request/billing reconciliation, not only self-consistent JSON. Likewise, a populated outcome reference does not prove that a reviewer checked the answer. The test establishes behavior of this supplied-record checker only.

Do not adapt the fixture into an SDK decoder by renaming keys. Streaming can carry cumulative or incremental observations; billing can arrive after response telemetry. A timeout with unknown usage stays unresolved until the owner obtains evidence or discloses an appropriately bounded estimate. Retain the attempt rather than declaring it free or silently excluding it.

7. Use a review record that survives a route change

The following filled example and blank counterpart have the same nine fields. Keep authoritative request and billing references in an approved evidence store, not in a broadly shared worksheet containing prompts, credentials or customer documents.

Filled comparison record

Question and owner
Illustrative E1: the application engineer compares cache accounting on the six-task cohort; the domain owner retains answer acceptance.
Population and cutoff
Q1–Q6, offsets 0/2/4/5/7/18 minutes, P1/P2 revisions; all supplied dispositions recorded after Q6.
Route and configuration
The documented Messages/Converse pair above; text-only, non-streaming, one five-minute checkpoint; actual invocation NOT EXECUTED.
Prior exposure and pacing
No real prior-cache evidence. Six sequential synthetic records; no warm-up cost or concurrency measurement claimed.
Usage and rate basis
Required fields supplied explicitly; fractional arithmetic at 2/3/0.25/8 fictional cost units per million tokens; no billed-cost reconciliation.
Outcomes and totals
Five stipulated acceptances and one rejection per arm; 0.089088 versus 0.111616 units. Q5 cost remains included.
Uncertainty and exclusions
Q3 cause unknown; no raw provider records, live outputs, latency, labor or full-workflow expenses.
Stop and recovery
Hold any live comparison on missing usage, broken joins, changed scope or failed quality. Preserve evidence and use the prior admitted route only within its existing authority.
Decision and next action
Accounting exercise only. Request independent review of the record and design a separately authorized, budgeted synthetic trial; no route selection follows.

Blank comparison record

Question and owner
Record the specific changed factor, accountable engineer and domain acceptance owner.
Population and cutoff
Attach ordered task IDs, arrival offsets, prefix revisions, exclusions and outcome cutoff.
Route and configuration
Record exact model, endpoint/API, source Region/profile, permitted handling, checkpoint/TTL and request-builder revision.
Prior exposure and pacing
Record known prior traffic, unknown cache state, deliberate warm-ups, concurrency and timing evidence.
Usage and rate basis
Attach native counters, adapter rules, unknowns, actual rate source/date/unit/currency and reconciliation owner.
Outcomes and totals
Record accepted/rejected/pending task identities, all attempts, category costs and the explicitly scoped denominator.
Uncertainty and exclusions
State unresolved usage, unmatched charges, sample limits, causal uncertainty and omitted workflow costs.
Stop and recovery
Name the stop owner, conditions, permitted prior route, in-flight reconciliation and restart evidence.
Decision and next action
Record hold or bounded recommendation, reviewer, remaining evidence and the next separately authorized action.

8. Choose what to test next, not a winner from a hit percentage

Request-level hits and token-level reuse answer different questions. Three requests with some cached input do not mean half of every prompt was reused. Retain category volumes and prefix sizes as well as the count of read-bearing requests. Report first uses, revised prefixes, quiet intervals and retries separately when those slices explain the production workload.

A one-hour checkpoint might suit longer gaps, but changes the write-price basis and retention review. Test it as a different configuration; do not apply its reuse assumption with the five-minute rate. A shorter prompt may remove useful repeated material or fall below the selected model's threshold. Compare its actual quality and cost rather than treating fewer tokens as an automatic improvement.

Do not tune the final task set until it generates the cache ratio you hoped to see. Improve a request builder on a development set, freeze it, then rerun representative arrivals under the same acceptance rule. Removing a volatile timestamp is a plausible implementation improvement if the task does not need it there. Reordering customer traffic merely to make a cache demonstration look good changes the operating question.

Before any live test, resolve the placement and retained-copy boundary. A cache lifetime is not a disposal receipt for every diagnostic or provider representation. Limit telemetry to permitted identities, counters and timing where possible. No cost result grants permission to share sensitive input or to widen a route's scope.

Start with one approved task family. Have another engineer inspect the ordered manifest, native counter mapping and one rejected task's retained cost. Hold the comparison if they cannot reproduce the boundary or if incomplete evidence could reverse the choice. If an authorized candidate later breaks quality or usage reconciliation, stop new experimental exposure, preserve in-flight records and restore only the previously admitted configuration under the release owner's control. Bring the resulting record to AI cost modeling or evaluation and observability when a scoped engineering review is useful. The worksheet and offline exercise do not require contact information.

Related services