Token Cost Is Not Cost per Completed AI Task
Measure AI workflow costs against verified completion, including retries, retrieval, human review and recovery. Compare changes without hiding failures.
Price the completed workflow, not just the model response
To measure cost per completed AI task, divide the attributable workflow cost by the number of distinct tasks that met a defined completion rule. Include unsuccessful attempts, retries, human review and recovery in the cost boundary. A low price per model call is useful engineering information, but it does not establish that the workflow is inexpensive to operate.
Define the outcome before collecting the cost. For a document-intake workflow, completion might mean that a supported record was accepted and the destination system confirmed the update. A generated extraction result, a successful HTTP response and a reviewer opening the item are intermediate events. None independently proves that the task is finished.
This article proposes a measurement method for engineering and operations teams. Its figures are hypothetical accounting inputs, not vendor prices, Ampity customer results or expected savings. The method complements token monitoring rather than replacing it. You still need call-level evidence to explain why the bill changed.
Choose one business unit and preserve its definition
Write the unit as a short contract: the item being counted, the required evidence, the relevant time boundary and the treatment of rejection or withdrawal. Keep that contract versioned. A change from “record accepted by the destination” to “draft produced” can improve the reported cost without improving the service.
Do not assume that every rejection is a failure. In an eligibility-checking workflow, a correctly supported rejection may be a completed decision. In a document-preparation service, a rejected draft may not satisfy the promised outcome. Choose the unit that answers the business question, then show the disposition categories separately so the denominator is not ambiguous.
Count a task once, even if it uses several models, retrieval requests or tool calls. Connect those events through a stable task identifier and retain attempt identifiers underneath it. Two calls that complete one task produce one business unit, not two. A duplicate submission may consume effort without representing a new requested outcome.
The FinOps Foundation's unit-economics capability distinguishes resource-efficiency measures from business-unit measures. Apply that distinction explicitly: cost per token explains a resource, while cost per accepted workflow outcome answers a different question. Neither metric should quietly inherit the other's denominator.
Build a cost boundary that can be reconciled
Start with the model charges, but do not stop there. List retrieval, embeddings, storage, tool services, orchestration, telemetry and the active human work needed to make the result usable. Include recovery and corrections where they belong to the measured workflow. Identify omissions rather than calling an incomplete numerator “total cost.”
Separate direct variable expenses from allocated shared costs. A dedicated extraction request can often be attributed to a task. A shared search index, platform subscription or on-call function may require an allocation policy. State the basis and report an unallocated balance if the evidence does not support a reliable split.
Human work needs a similarly explicit policy. Measured review minutes multiplied by an agreed loaded hourly cost can estimate labor consumption. It is not automatically a cash saving if fewer minutes are used next month: salaries, staffing commitments and other work may remain unchanged. Keep consumed effort, avoidable spend and available capacity as different measures.
Development and evaluation costs also matter, but mixing one-time implementation costs into a single pilot week can distort comparison with steady operation. Present recurring operation and a separately stated implementation allocation. Explain the period and useful-life assumption if you amortize a cost. This is operational planning, not a substitute for the organization's accounting policy.
Worked example: the bill is small, the review cost is not
Assume a synthetic intake cohort contains 1,000 distinct submissions. At the observation cutoff, 800 meet the chosen accepted-record completion rule, 100 are rejected and 100 remain unresolved. For this example, rejected items are reported separately and do not count as accepted-record completions. All attributable cohort costs remain in the numerator.
Suppose 240 items require active review, with five minutes of total handling per item across all visits. That is 1,200 minutes, or 20 hours. At an illustrative loaded rate of $30 per hour, review effort contributes $600. These inputs are chosen to make the arithmetic inspectable; they are not staffing advice.
| Cost category | Illustrative cohort cost | Boundary used here | | --- | --- | --- | | Model requests | $300 | Includes successful calls and billable repeat attempts | | Retrieval and orchestration | $120 | Directly attributed processing expenses | | Active human review | $600 | 20 hours multiplied by $30 per hour | | Recovery and correction effort | $80 | Additional effort, not included in review minutes | | Allocated shared operation | $100 | Agreed share of relevant platform costs | | Total within this boundary | $1,200 | Excludes one-time implementation in this example |
The model-only cost per submitted item is $300 divided by 1,000, or $0.30. The workflow cost per accepted completion at this cutoff is $1,200 divided by 800, or $1.50. Those figures answer different questions. Calling the first figure the cost of completing the task hides both the outcome definition and most of the measured expense.
The 100 unresolved items may incur more cost or produce more completions later. Mark the cohort result as provisional and update it under the same rule. Do not move their existing expense out of the ledger merely because their outcomes are inconvenient. If the denominator is zero, report that no cost-per-completion value can be calculated; do not substitute zero cost or divide by attempts.
Retries and review changes can reverse a model-price saving
Now consider an alternative configuration on a comparable synthetic cohort. Its model cost is $200, but its review effort rises to $900. Retrieval and orchestration stay at $120, recovery at $80 and shared operation at $100. Total measured cost is $1,400, and only 750 tasks meet the same completion rule by the comparable cutoff.
Its cost per accepted completion is approximately $1.87. The model component is cheaper, yet the measured workflow is more expensive per accepted outcome. That does not prove that cheaper models always create more review. It shows why the team must test the whole process rather than infer workflow economics from a provider's unit price.
Record why repeat work occurred. A transient processing retry, a schema repair, a user clarification and reconciliation of an uncertain business write are different events. Some are necessary. Others reveal a preventable defect. They can have different charges and different authority requirements, so avoid one undifferentiated “retry cost” bucket.
Reduce the cause of costly rework before deciding to remove a control. Better evidence routing may reduce unnecessary reviewer visits. A narrower task contract may prevent repeated output repair. Skipping required review can make the spreadsheet look better while increasing unsupported decisions. Cost optimization needs a quality boundary, not only a smaller numerator.
Reconcile provider usage without treating it as business completion
Use provider records to validate the model-cost component. Preserve the billing period, workspace or project, relevant pricing agreement and the categories actually charged. Check whether cached input, tools or other services are included. Do not assume every token has the same rate or that every request has a nonzero charge.
Anthropic's Usage and Cost API documentation provides organization-level reporting and describes product-specific access and coverage limits. The important application boundary remains separate: a provider's usage report cannot confirm whether your destination record was accepted or the reader's problem was resolved.
Join internal task and attempt records to usage where the available identifiers support it. If the provider exposes only aggregated cost, use a stated allocation method and retain the reconciliation difference. Do not claim exact per-task invoiced cost from token estimates alone. Estimates can be useful when clearly labeled and checked against the actual billing totals.
Keep cost instrumentation proportionate. A task identifier, model revision, usage categories, timing, outcome and exception reason may support the investigation without storing the entire submitted document or conversation. Financial observability should not become an unnecessary sensitive-content archive.
Compare cohorts and periods without mixing unfinished work
A cohort view follows submissions that began in a defined window and their eventual costs and outcomes. A period view measures operating expenses and completions during that period, including work that began earlier. Both can be useful, but they cannot be swapped without explanation. A backlog drain may make period cost per completion look unusually good.
For a controlled comparison, use the same task definition and a comparable case mix. Record document complexity, source quality, review policy, service objectives and observation maturity. A configuration tested on easier cases may appear more economical without being better for the actual workload. Stratify important differences rather than hiding them in one average.
Show pending age and remaining work next to the ratio. A low cost per accepted item can coexist with a growing queue of difficult cases. Report acceptance, correct rejection, withdrawal, unresolved work and reopened decisions separately. That keeps the number connected to the service people actually receive.
Do not infer return on investment from this operating metric alone. Revenue, avoided loss, cycle time, implementation cost and the alternative workflow may matter to the commercial decision. Name the measured scope and leave unsupported benefits out. A clear narrow measurement is more useful than a broad savings claim built on missing evidence.
Measurement limitations: a ratio cannot settle the decision
Do not use a single cost-per-completion ratio to compare unrelated outcomes. Preparing a draft and executing a verified record change are different services, even if the interface calls both a task. A blended average can also conceal a high-cost task family that needs separate review.
The metric becomes weak when labor measurements exclude hard cases, shared costs have no defensible allocation or the outcome sample is too immature. Record those gaps and show a sensitivity range rather than disguising estimates as invoiced precision. Cost also cannot compensate for an unauthorized action: a cheap completion outside the approved scope is not an acceptable result. Use independent quality and authority tests alongside the economic comparison.
Use a change record before optimizing production
Start with one task family and a reconciled baseline. Record the completion contract, cost categories, data sources, unresolved allocation, review effort and quality checks. Have the workflow owner and the person responsible for cost reporting agree the definition before testing a candidate change.
Change one consequential factor at a time where feasible. Compare the measured total alongside completion, error categories, waiting time and rework. A caching experiment needs freshness and permission checks; a model-routing experiment needs relevant acceptance tests; a retry change needs evidence that it does not duplicate business effects. The cheaper path must still satisfy the task contract.
Set the stop condition in advance. Pause expansion if the candidate breaches the agreed quality or authority boundary, if cost attribution cannot be reconciled, or if unfinished work makes the comparison misleading. A pilot can end with insufficient evidence. It does not need to produce a favorable savings percentage to be useful.
Read review-queue capacity for the human-work bottleneck and tool results versus business completion for the completion boundary. The AI change-release playbook provides a way to evaluate behavior changes before expanding exposure. For help implementing measurement and controls, explore production AI systems or share your workflow question. Resource access does not require a contact request.