Measure AI Cost per Completed Task
Build a reconciled task-cohort ledger that includes retries, review and recovery. Define accepted completion, disclose pending work and test the economics before scaling.
trigger="Model spend is visible, but the team cannot reconcile it with accepted work, retries, reviewer effort or pending tasks." owner="The application owner accountable for the workflow's accepted output and operating budget." participants={['Application engineer', 'Domain acceptance owner', 'FinOps or finance reviewer', 'Review-team lead', 'Operations owner']} prerequisites={['One defined business task and its completion contract', 'An approved observation population and cutoff', 'Provider billing evidence and application task records', 'A documented labor basis and shared-cost allocation policy']} outputs={['A versioned completion and cohort definition', 'Reconciled task, attempt and cost records', 'A cost-per-completion worksheet with sensitivities', 'An owned optimization experiment or a documented measurement hold']} doneWhen={['Retries cannot inflate the completion count', 'All task dispositions retain their attributable cost', 'Unmatched charges and pending work are disclosed', 'Another reviewer can reproduce the ratio and its cost basis']} />
Measure accepted work, not just model responses
Use this playbook when you need to decide whether an AI-assisted workflow is economical enough to operate or expand. The output is a reproducible measurement record, not a provider-price comparison. It connects the cost of a defined task population to the work that actually met the completion contract.
The procedure is proposed guidance. Its dollar values and task counts are synthetic and do not represent Ampity client outcomes or current vendor prices. Use your own approved data, actual billing basis and domain acceptance rules. A cheaper ratio is not an improvement if the task became easier, the output became less useful or the system stopped counting difficult requests.
Start with the cost-per-completed-task article for the core distinction. Here, you will define the denominator, link attempts to tasks, reconcile the numerator and test a specific optimization. Stop the measurement when its evidence is incomplete rather than presenting a precise number with an uncertain meaning.
1. Define the business unit and acceptance contract
Owner: application owner with the domain acceptance owner. Output: a versioned unit definition. Name one task that produces a useful result. Examples include a reviewed intake record or a supported answer to a defined question. Do not combine a lookup, a complex investigation and an external write into one undifferentiated completion count.
State what qualifies as accepted completion, who decides and where that decision is recorded. A valid JSON response, a provider success code or an agent's assertion is not automatically a completed business task. Include any required downstream state and the treatment of later corrections or reversals.
The FinOps Foundation's unit-economics capability connects technology cost to meaningful business units and emphasizes consistent metric definitions and assumptions. Apply that principle with your domain owner. Do not copy a generic unit merely because it is easy to extract from a dashboard.
Keep quality and timeliness requirements visible. If an answer is correct but arrives after the useful deadline, decide whether it belongs in the accepted population. Version that rule before comparing alternatives. Otherwise a change in acceptance policy can look like a technical cost improvement.
2. Freeze a cohort and an observation cutoff
Owner: analyst or application engineer. Output: a cohort manifest. Choose a population by intake time, task family and eligibility rules. Retain its stable task identifiers. Then record how long outcomes are observed and which costs are included through that cutoff. Use one timezone convention and explicit boundaries.
Do not mix the month's charges with unrelated tasks that happened to finish that month unless that is the deliberately defined period metric. A cohort measurement follows the selected tasks, including their later attempts and review. A period measurement answers a different operating question. Label both if you need both.
Keep pending tasks in the population and disclose them. They are not accepted completions, but their accrued cost does not disappear. If pending work is substantial, call the ratio provisional and show the pending count and age. Set the follow-up observation date rather than silently dropping those records.
Record exclusions and their rationale. Sandbox traffic, load tests and other task families may belong in separate ledgers, not be ignored. The reviewer needs to see where excluded charges went. A cohort filter is an accounting decision and should not be edited simply to make the result attractive.
3. Link attempts and effects to stable task identities
Owner: application engineer. Output: task and attempt joins. Assign one business-task identity with separate identifiers for generation attempts, retrieval work, tool operations and review events. A retry, fallback or correction creates another attempt, not another accepted task. Retain the original identity through queueing and recovery.
Choose the authoritative completion record. If the task updates another system, use the domain state or appropriate readback rather than an assistant's final sentence. Resolve duplicate events by identity and revision. Do not count a repeated completion notification twice because it arrived through two channels.
Inspect a sample of joins manually. Follow one straightforward completion, one rejected item, one pending item and one retried task from intake to outcome. Verify that the records describe the same business intent. Stop if identifiers are reused or if the tracing system cannot distinguish attempts from tasks.
Minimize measurement data. Cost investigation generally needs identifiers, durations, configurations and dispositions, not unrestricted copies of every document or conversation. Protect the ledger and retain payload access only where a specific investigation requires it under an approved policy.
4. Collect model and service charges on a stated basis
Owner: FinOps reviewer with the application engineer. Output: provider-charge evidence. Retrieve the available usage and cost reports for the actual account and product. Capture the observation window, reporting granularity, currency and whether values are estimated, accrued or billed. Keep late-arriving adjustments visible.
Claude's Usage and Cost API documentation describes organization reporting and product-specific access requirements. OpenAI's organization usage reference separates usage and cost resources. Neither source establishes your application's accepted completion count; that must come from the task records.
Use provider evidence to reconcile application estimates, not to manufacture task-level precision the reporting granularity cannot support. If a report aggregates several workflows, allocate it with a documented method or keep an unallocated remainder. Do not assume every aggregated charge can be joined to a request identifier.
Avoid counting the same charge twice through both provider billing and a cloud invoice that already includes it. Identify reseller or marketplace billing paths, credits and committed-spend treatment with finance. This playbook does not prescribe tax or financial-accounting treatment; it requires a clear basis for the operating comparison.
5. Include review and recovery effort
Owner: review-team lead and operations owner. Output: effort ledger. Record reviewer handling time, correction effort and incident recovery attributable to the selected tasks. Distinguish active handling from queue waiting. Waiting affects service performance, but multiplying every elapsed hour by a labor rate can overstate direct effort.
Define the labor basis with finance. It might be loaded internal effort, incremental contractor expense or another approved costing method. Identify estimates and sampling uncertainty. If review is sampled, explain how the sample covers task complexity and how the estimate expands to the cohort.
Include failed and rejected tasks. Their model calls and review effort were still consumed. Recovery can include reconciliation of uncertain writes and correction of accepted records. Avoid adding both a total incident allocation and the same recovery minutes again as direct labor.
Time saved is not automatically cash saved. Report released capacity separately from reduced expenditure unless staffing or contracted charges actually changed. An operating-cost worksheet can support a decision about capacity without claiming a financial benefit that the organization has not realized.
6. Allocate shared operations without hiding the remainder
Owner: FinOps or finance reviewer. Output: allocation-policy revision. Identify shared retrieval storage, runtime infrastructure, observability and operational support included in the numerator. Separate directly attributable charges from pooled costs. Choose a distribution basis that reflects the intended management question.
The FinOps allocation capability describes direct and shared allocation strategies and the need to document apportionment. Your worksheet should name its selected rule, the total pool and how the shares reconcile. An equal split and a measured-usage split can produce different results without either being a vendor-recommended universal choice.
Show unallocated costs explicitly. If ownership is missing, leave a visible remainder with a remediation owner instead of distributing it arbitrarily to make reconciliation appear complete. Decide whether the current evidence is sufficient for a provisional analysis or whether the comparison must be held.
Keep setup and recurring costs distinct. One-time integration effort can matter to the investment decision but should not silently enter an ordinary operating ratio. If you amortize it for a sensitivity case, show the horizon and amount separately. Do not confuse the steady-state worksheet with a full investment-return calculation.
7. Reconcile the ledger before dividing
Owner: analyst with an independent finance reviewer. Output: reconciliation record. Match provider totals, application estimates and allocated cost categories. Explain differences from reporting delay, currency conversion, missing joins or excluded populations. A missing match is a finding, not a zero-cost task.
Use the following worksheet fields as a starting point. Keep source references and policy revisions alongside the values so another reviewer can reproduce the result without having to infer how the spreadsheet was constructed.
| Field | Required record | Verification | | --- | --- | --- | | Cohort | Task family, IDs and intake boundaries | Distinct task count matches manifest | | Outcome | Accepted, rejected, cancelled or pending at cutoff | One disposition per task under the contract | | Direct cost | Model, retrieval and other task-attributed charges | No duplicate billing path | | Effort | Active review and recovery time with costing basis | No overlap with shared incident allocation | | Shared cost | Pool, distribution rule and allocated share | Shares plus remainder match the pool | | Adjustments | Credits, late charges and estimation differences | Source and treatment are documented | | Result status | Final, provisional or held | Pending and unmatched amounts remain visible |
Set an acceptable reconciliation tolerance before interpreting the result. The tolerance belongs to the owner and use case, not to this article. Small differences may be immaterial to one decision and decisive to another. If uncertainty could reverse the proposed choice, hold the conclusion and improve the evidence.
8. Calculate the ratio and disclose its denominator
Owner: analyst. Output: reproducible calculation. Divide the reconciled cohort cost by distinct accepted completed tasks at the observation cutoff. Include costs from rejected, cancelled and pending cohort tasks in that numerator when they fall within the defined basis. Disclose each disposition count separately.
For a synthetic cohort of 1,000 tasks, suppose 800 are accepted completions, 100 are rejected and 100 remain pending. Costs are $300 model charges, $120 retrieval, $600 review effort, $80 recovery and $100 shared operations. Total cost is $1,200, and the provisional ratio is $1.50 per accepted completion.
The $600 review assumption represents 240 reviews at five active minutes each: 1,200 minutes, or 20 hours, at a synthetic $30 hourly effort rate. It is not a vendor price or a claim of realized labor savings. The model-only figure of $300 divided by 1,000 submissions is $0.30, but it answers a different question.
If no task completed, do not return zero or divide by a substituted denominator. Report no accepted completions and the cost consumed, with the reason. A finite cost-per-completion ratio is unavailable. Keep pending work visible and reassess at the agreed cutoff rather than implying that a failed workflow is free.
9. Test the drivers that could change the decision
Owner: analyst with the domain owner. Output: sensitivity worksheet. Vary review frequency, handling time, completion rate and relevant service costs. State which values are observed and which are scenarios. Do not represent a sensitivity case as a forecast unless its demand and behavior assumptions are supported.
In the synthetic example, doubling active review time to ten minutes increases review effort from $600 to $1,200. With the other costs and 800 completions unchanged, total cost becomes $1,800 and the ratio becomes $2.25. If only 600 tasks completed instead, the original $1,200 numerator produces $2.00 per completion.
Neither change should be assessed in isolation from quality and queue performance. Reduced review can lower measured effort while increasing errors or later corrections. Higher completion counts can reflect a relaxed contract. Require the same accepted-output definition when comparing cases intended to show an improvement.
Segment by meaningful task complexity where the population warrants it. Averages can hide expensive exceptions and a changing input mix. Keep segment definitions stable and report their volumes. A lower overall ratio caused only by more easy tasks should not be described as an engineering optimization.
10. Choose one bounded optimization experiment
Owner: application owner with the release owner. Output: experiment contract. Select one driver supported by the ledger: repeated generation, unnecessary retrieval, excessive review time or unresolved recovery effort. Define the change, expected mechanism, task population and observation window before altering the system.
Measure accepted completions, quality, latency, review demand and cost together. Compare like-for-like task populations or document differences explicitly. Retain baseline configuration and evidence. A smaller model bill with fewer accepted outcomes is not enough to approve the change.
Specify stop conditions for unauthorized effects, unacceptable quality, queue growth or measurement failure. Use an isolated test first where the change can cause business effects. The experiment owner must be able to disable exposure without losing the records needed to reconcile work already in progress.
Keep this procedure separate from deployment authorization. A favorable worksheet supports a recommendation, not automatic release. Use the AI change-release playbook to test behavior and rollback before expanding the proposed optimization.
11. Recover when the measurement fails
Owner: analyst and operations owner. Output: corrected or held report. If task joins break, charges arrive late or the completion definition changes, stop publishing the affected comparison. Preserve the last verified report and label the current one provisional or held. Do not replace missing evidence with a zero or a guessed completion.
Rebuild the affected window from retained source records under the same policy, or publish a clearly versioned restatement. Explain which numerator or denominator changed and why. Keep the original result available to authorized reviewers so the correction does not erase the investigation trail.
If an optimization causes harm, stop it using the defined release control and reconcile affected tasks. Do not rerun every pending action merely to improve the completion count. Recovery must respect current permissions, task state and duplicate-effect safeguards. The economic measurement follows the recovered truth; it does not authorize recovery actions.
Document limits. A small pilot, sampled labor or immature pending cohort may support further measurement rather than a scale decision. The correct conclusion can be that the current evidence cannot distinguish the alternatives. That is more useful than a precise ratio that cannot survive a basic reconciliation review.
12. Close with an owned decision and reusable checklist
Owner: application owner. Output: accepted measurement record and next action. Review the result with engineering, the domain owner and finance. State whether it supports an experiment, continued operation, a hold or additional instrumentation. Name the next observation date and the person responsible for resolving uncertainty.
Before closing, confirm the acceptance contract is versioned, the cohort is reproducible, retries are not extra tasks, all dispositions retain cost, provider evidence is reconciled, shared allocation is visible and review effort uses a stated basis. Confirm that pending work, unmatched amounts and excluded costs are disclosed alongside the ratio.
Use this review checklist with an evidence reference and owner for each item:
- Completion is independently supported by the versioned domain contract.
- Cohort intake boundaries and outcome cutoff can be reproduced.
- Each attempt joins to one task, and each task has one disposition at cutoff.
- Direct charges and shared allocations reconcile without duplicate billing paths.
- Review and recovery effort do not overlap with another cost category.
- Pending tasks, unallocated amounts and estimates remain visible.
- Sensitivity cases preserve the acceptance contract and distinguish assumptions from observations.
- The next experiment has stop conditions, a recovery owner and separate release approval.
Have an independent reviewer reproduce the synthetic arithmetic and one actual cohort calculation from the retained evidence. Record any disagreements about completion or allocation. Resolve them through the policy owner rather than choosing the treatment that creates the lowest reported cost.
For queue constraints, read review capacity in AI automation. For uncertain effects, use the action-recovery playbook. If you want help applying this measurement, explore production AI systems or share your measurement question. Downloading or reading the resource does not require contact information; an enquiry is a separate optional choice.