Compare Bedrock and Direct AI APIs Without Losing the Tools Your Team Needs
Compare Amazon Bedrock and direct AI APIs using task acceptance, retained employee tools, contract obligations, transition costs and controlled migration gates.
You can evaluate an application capability on Amazon Bedrock while keeping the ChatGPT or Claude subscriptions people still use. Start with the work and the contracts attached to it. An API migration does not establish that a chat subscription is redundant, that stored state transfers, or that the whole workflow costs less.
This playbook helps an engineering lead and finance owner produce a task, dependency and spend comparison before changing traffic or cancelling a contract. It covers direct application APIs versus an exact Bedrock route, retained employee tools, trial costs, parallel operation and a controlled handover. It does not select a universally better model, determine legal compliance or authorize paid tests. The worked company, prices, counts and outcomes are fictional teaching inputs. No AWS account read, model invocation, customer evaluation, subscription cancellation or deployment occurred.
Follow the decisions separately
This process view shows responsibility and evidence, not network architecture. Moving an application capability does not by itself permit removing an employee product or deleting source-held state.
Capability, production release, full costs and contract retirement have different evidence and owners. The branches are not a sequence of automatic approvals. Employee-tool replacement remains a separately accepted decision.
| Decision point | Evidence going in | Possible result | What it does not authorize |
|---|---|---|---|
| Keep or replace employee tools | Required staff tasks, plan benefits and contract terms | Retain or investigate a separately accepted replacement | Application API migration |
| Trial an application capability | Complete task inputs, exact supported route, access and processing permission | Bounded trial, retained route or hold | Sending unapproved data or removing a source dependency |
| Adopt a split | Acceptance, failures, operating evidence and full horizon costs | Retain, split or hold | Automatically claiming savings or cancelling contracts |
| Retire a contract or object | Continuing consumers, replacement proof, obligations, recovery and effective date | Separately authorized retirement with readback, or retain | Deleting unrelated state or cancelling useful employee seats |
1. Put the right people around one decision
The trigger is a proposal to consolidate AI spend, move an application to Bedrock or stop paying a direct provider. The sponsor should name one application capability and the decision they will make with the result: retain, trial, split, migrate or stop the proposal. Avoid beginning with a target percentage saving before anyone has described the work.
The application owner records task and state dependencies. Finance supplies invoices, commitments and renewal dates. The data/security owners decide permitted inputs and processing boundaries. A product reviewer defines acceptable outputs. The operating owner sets failure behavior, observation and rollback. Procurement or the authorized contract owner decides whether a cancellation can take effect. One person may hold several roles, but each decision still needs an accountable owner.
Before collection, agree where sanitized evidence will live, who can inspect it and which activities are permitted. Reading a bill, collecting usage and invoking a model are different permissions. A synthetic evaluation can incur usage charges. Do not grant general administration or send private customer data merely to complete this worksheet.
Expected output: a dated comparison boundary and a named next decision. Stop if nobody owns task acceptance, if required evidence cannot be obtained safely or if the proposal is already presented as a completed saving.
2. Separate employee products from application APIs
OpenAI's billing guidance identifies separate ChatGPT and API billing systems. Claude's current billing guidance likewise distinguishes paid Claude products and Console API access, while noting plan-specific monthly API credits for Max and Team. Verify the actual plan and credit terms rather than assuming every subscription includes no API benefit.
The finance owner creates one row per contract or independently billed product. Record the purchasing entity, product, account/workspace scope, payer, purpose, plan, metering, commitment, renewal/cancellation terms and evidence date. Store an internal evidence reference rather than a complete invoice or personal billing identity in a broadly shared worksheet.
| Purchase category | Work to inventory | Evidence needed before removing spend |
|---|---|---|
| Human-facing chat/workspace subscription | Staff drafting, analysis, shared workflows and any dependent integrations | Named users, actual required features, permitted alternative, renewal terms and accepted replacement evidence |
| Direct application API | Requests from the product, hosted tools, stored state and supporting storage | Calling applications, credentials ownership, exact operations, continuing dependencies and remaining commitments |
| Bedrock candidate usage | Selected model/API route and related AWS services | Applicable account access, route support, processing approval, usage dimensions and dated price basis |
| Adjacent service or commitment | Retrieval, storage, support, reserved capacity or an external tool | Who still consumes it, whether it is cancellable, and when charges actually change |
A shared email address is not evidence that two charges buy the same function. A seat may remain useful while an application API moves. Conversely, a seat may be independently unnecessary even when the application stays direct. Keep those decisions separate so a model trial cannot accidentally remove a team's everyday tool.
Gate: every proposed removed charge has a named consumer and contract owner. Unknown purpose means retain pending investigation, not zero cost or permission to cancel.
3. Inventory complete tasks before assigning routes
The application owner samples the real operations the product promises. A “support assistant” can contain self-contained passage summaries, hosted document answers and continuing conversations. Treat each as a separate task contract until evidence supports combining them.
For each capability, record authoritative inputs, assembled context, acceptable output, deadline, permitted effects, external tools and state ownership. Identify whether a request needs a provider-held conversation, uploaded file, vector store, cached prefix or tool result. OpenAI's conversation-state documentation describes durable conversation objects and stored items. Its file-search documentation describes the hosted retrieval path and vector-store dependencies. These are dependencies to inspect, not portable objects merely because another API accepts text.
Use the stored-state inventory for object-specific treatment. Keep a capability-level routing record: eligible task, retained dependency, exact route and held uncertain attempt. This playbook adds the contract/spend disposition, rather than designing an adapter.
The first reusable artifact is a task-contract ledger:
| Field | What the owner records |
|---|---|
| Capability and user | Named operation, revision, original task identity and who relies on it |
| Accepted result | Required evidence, retained conditions, deadline and prohibited effects |
| Dependency | Exact input/state/tool/service, authoritative revision and continuing consumers |
| Current route | Provider, account/project reference, endpoint, API and model/version |
| Candidate route | Complete target tuple or UNKNOWN; selected feature settings |
| Contract disposition | Retain, investigate replacement or separately retire; earliest effective date |
| Missing evidence | Named owner, bounded collection action and stop condition |
Gate: a self-contained candidate has all inputs needed without borrowing inaccessible source state. A composite task remains composite even if one inference step can move.
4. Bind the candidate to a supported route
AWS's API compatibility reference separates runtime API families and warns that aggregate compatibility can span different endpoints. A model name or a familiar SDK method cannot establish the chosen endpoint's feature support. Keep provider, model/version, endpoint family, operation, source Region, inference profile where applicable, account/caller and selected features together.
A documentary candidate for a bounded text trial is Claude Haiku 4.5 through Bedrock Runtime Converse, profile us.anthropic.claude-haiku-4-5-20251001-v1:0, originating from us-east-1. The model card documents this Runtime route as cross-Region rather than in-Region. Its support must be rechecked at the trial date. This is a route to investigate, not a recommended model or observed account entitlement.
The Converse request contract identifies its model and message inputs. Do not send a source provider's conversation identifier and assume it becomes a continued conversation. A different API/model combination needs its own record, even when the visible product stays unchanged.
The account/security owners then collect access and placement evidence. AWS's model-access guidance describes applicable activation/licensing prerequisites and distinguishes endpoint-specific requirements. A successful form submission can leave an agreement pending. A documentation row or administrator demo does not establish the application's caller permissions or usable quota.
Processing and retained-copy approval belongs in the existing Bedrock placement paper. Resolve that boundary before exposure. Do not broaden permissions, infer India-only processing from an Indian source endpoint or quietly substitute another route to complete the commercial comparison.
Output: exact candidate binding with documented, observed and unknown facts separated. Hold the trial on unsupported required features, missing authority or an incompatible processing boundary.
5. Test the application contract, including failures
The product reviewer defines expectations before either candidate produces answers. Start with representative permitted tasks, then include cases whose failure would change the decision. For a summary, an exception in the source must remain visible. For an answer, the source must support it and be available to the user. For a proposed action, the application must enforce current authority independently of the model.
Record original task IDs, input revisions, task families, evaluation split and reviewers. Keep variants of one document together when splitting data. Count a task once after retry or human correction, using the same acceptance definition for both candidates. Preserve held tasks, failed attempts and unresolved outcomes rather than dropping them from the denominator.
| Proposed check | Expected evidence and failure disposition |
|---|---|
| Source exception survives a summary | Reviewed answer against the original passage; omission fails acceptance |
| Source-held state is required | Candidate remains on the retained route or holds; no fabricated continuation |
| Caller is denied or quota is exhausted | Owned user-visible hold/degradation; no permission widening or unreviewed route substitution |
| Timeout leaves an attempt uncertain | Attempt recorded and reconciled under the recovery policy; no automatic duplicate business effect |
| Model returns an incomplete response | Release is withheld under the exact response contract |
| Model proposes a tool | Proposal, argument validation, current authority and actual execution receipt remain distinct |
Streaming, images and tools need their specific contract checks; the existing Bedrock portability articles own those mechanisms. An initial text-only trial should not inherit their evidence. Keep safety, output quality, latency/deadline and cost as separate results. A cheaper candidate that fails a mandatory dimension remains unsuitable for the task.
Gate: an independently reviewed bounded evaluation packet supports the proposed task, not the entire provider catalogue. These are proposed checks only; none of the model tests above were executed.
6. Reconcile costs that can actually change
Finance uses one horizon and currency, with disjoint cost rows. Bedrock pricing varies by model, modality and operating tier and lists separate service charges. Enter the applicable units and dated sources; do not turn cache, batch, storage or tool charges into a single assumed token rate.
Separate four views: current application cost, retained employee-product spend, transition cost and resulting portfolio cash outflow. For each row distinguish observed invoice/usage, documented rate and an assumption. A missing value remains unknown. A known zero needs a basis. Do not allocate the same storage bill to retrieval and operations, or count review hours twice.
Contracted spend may remain payable after traffic stops. Record cancellation notice, effective date and minimum commitment. Sunk cost, remaining obligation and avoidable future outflow answer different questions. A reduction in measured requests need not reduce the invoice if a commitment still applies. A plan credit may alter a scoped charge without replacing the plan's fee.
Include implementation, evaluation, retained provider usage/state, overlapping bills, added operations, reviewer effort and applicable taxes/FX where they affect the actual decision. Report exclusions. Show a no-credit view first; any separately confirmed applicable credit needs balance, eligible charge, expiry and contract evidence. This worksheet does not infer an AWS award or entitlement.
Use the AI cost comparison tool in the mode appropriate to the question. The accepted-result mode compares ordinary token, retrieval, review and operating inputs. The route total-cost mode adds dated monthly charges, retained seats, commitments, transition costs and separately scoped confirmed credits. Both require an explicit task boundary and evidence; neither establishes feature equivalence, account eligibility or contract cancellation rights. Unsupported meters need their own reconciled lines, not an invented conversion. Keep employee seat cost outside an application-task denominator when the seats serve other work.
7. Work a comparison that changes after transition costs
The fictional team has 20,000 original application tasks each month. A proposed Bedrock split moves 12,000 self-contained summaries; 8,000 tasks retain direct-provider features. Employee tools remain separately useful and cost a stipulated USD360 per month. These are teaching amounts, not provider prices, invoice observations or an Ampity case.
The example assumes repeated monthly workload/cost amounts for six months, no inflation, credits, tax, FX, financing or reserved commitment. The owner stipulates that both compared application scopes preserve the promised task capability. Actual field evidence for that premise remains absent. No spend or quality result can authorize the missing trial.
| Monthly included cost, USD | Stay direct | Proposed split | Downside split |
|---|---|---|---|
| Summary inference, all included attempts | 900 | 450 | 700 |
| Retained-feature inference | 300 | 300 | 300 |
| Retained supporting storage | 100 | 100 | 100 |
| Application review effort | 1,200 | 1,300 | 1,600 |
| Other disjoint application operations | 400 | 550 | 650 |
| Application subtotal | 2,900 | 2,700 | 3,350 |
| Retained employee products | 360 | 360 | 360 |
| Portfolio monthly total | 3,260 | 3,060 | 3,710 |
For the normal split, stipulated accepted original application tasks are 19,000; the baseline has 19,200. The minimum is stipulated at 95 percent, or 19,000 of 20,000, with a separately stipulated passed safety/deadline status in both scenarios. The downside has 18,700 accepted tasks and fails the threshold. These are synthetic labels, not observed evaluations or universally suitable acceptance criteria.
The application unit costs are 2900 / 19200 = USD0.1510416667 baseline and 2700 / 19000 = USD0.1421052632 split. Employee-tool costs are excluded from these ratios because their independent tasks were not counted. The lower split ratio does not erase the baseline's 200 additional accepted tasks or prove comparable value across different failure cases. Review which tasks failed before making a business choice.
Migration and evaluation build effort costs a stipulated USD6,000 once. An additional transition-only dual-run allocation is USD800. That allocation excludes ongoing application rows already counted; it is not the entire old route bill counted again. Thus transition cost is USD6,800.
Over six months, staying costs 6 × 3260 = USD19,560. The split costs 6 × 3060 + 6800 = USD25,160, a USD5,600 increase. Monthly recurring outflow is USD200 lower, yet the six-month proposal costs more. With unchanged assumptions, no additional charges and the same monthly difference continuing, the simple transition payback is 6800 / 200 = 34 months. This conditional arithmetic extends beyond the six-month decision horizon, so it is not an accepted forecast.
The downside split costs 6 × 3710 + 6800 = USD29,060, USD9,500 above staying, and fails the stipulated task threshold. Its incurred application unit cost remains 3350 / 18700, about USD0.1791443850; it cannot be presented as a suitable winner. If accepted count were zero, the ratio would be undefined while costs remained incurred.
A procurement owner who removes USD360 of seat spend from the split has changed an unrelated task assumption. The record must identify which staff work disappeared or which approved replacement covers it. Without that evidence, the removal is rejected. If an actual seat reduction is justified later, record its notice and effective date as a separate scenario rather than backdating it into all six months.
8. Choose a bounded dual-run, not unlimited duplicate traffic
The application owner defines which traffic may be duplicated and why. Shadow evaluation, human-reviewed assistance and customer-visible replacement have different exposures. Use fabricated or explicitly permitted inputs first. Do not mirror external writes, payments or communications just because the target request looks harmless. A model-generated tool proposal remains subject to the destination's controls.
Name the trial duration, task population, spend cap, observers and stop owner. Include old-route and candidate charges in the same window. Observe accepted outputs, held work, deadline failures, additional attempts and reviewer demand. A budget alert is an observation, not automatically an enforced cap; the operating owner must know the actual stop mechanism and outstanding work after stopping.
Stop on changed model/API/profile/caller binding, unknown input classification, unexpected retained content, missing receipts, unsafe proposed effects or reviewer backlog beyond the agreed operating bound. Preserve useful retained functions rather than disabling the entire product. Choose a smaller trial when full parallel operation is too expensive or unsafe. Choose the experiment mode before collecting transition and spending evidence. A shadow run that cannot affect customers and a customer-visible replacement answer different questions.
Gate: trial evidence reconciles the admitted population, costs and outcomes. Unknown usage or unsettled attempts remain open items, not clean savings.
9. Hand over the split before retiring a dependency
The operating owner receives the versioned route/configuration references, release checks, cost observations, queue ownership, incidents and recovery instructions. A rollback defines what happens to new work separately from in-flight tasks. Preserve their original attempt identities and state. A late response must not be attached to a new task after changing routes.
Run proposed adverse rehearsals under separately authorized scope: candidate unavailable; retained provider unavailable; changed state reference; delayed output; missing cost receipt; and a task misclassified as eligible for the new route. Define expected user behavior and which owner investigates. Do not use a healthy second provider as proof that replay is permitted.
Before cancelling any direct contract or deleting provider objects, check continuing consumers, export/reconstruction needs, unresolved attempts, retention obligations, backups and required access during recovery. Cancellation and deletion are separate actions. A stopped API bill does not establish disposal of all previously created copies, and returning the routing configuration does not restore deleted state.
The procurement owner uses a cancellation packet containing the exact contract, removable scope, continuing consumers, replacement evidence, earliest effective date, remaining obligations, authorized actor and resulting-state readback. A named owner approves the action under company procedure. This article supplies no cancellation API, credential revocation command or destructive cleanup instruction.
10. Preserve an inspectable decision packet
Keep the decision readable enough to challenge without reconstructing the whole trial. Store authoritative records in their permitted locations and share sanitized references. The second reusable artifact is the commercial disposition record below. Every blank or unsupported field stays UNKNOWN until its owner supplies evidence.
| Record field | Filled fictional example | Blank record prompt |
|---|---|---|
| Decision/window | C-PB1; six months; USD | Sponsor, revision, horizon and currency |
| Capability boundary | 12,000 summaries eligible; 8,000 retained-feature tasks/month | Original tasks, input/state/tool dependencies and exclusions |
| Employee-product scope | USD360/month retained; different staff tasks | Users, required features, plan benefits and proposed removals |
| Exact candidate | Documentary Haiku Runtime Converse US profile above; account evidence UNKNOWN | Model/version, endpoint/API, source/profile, caller, feature settings and check date |
| Acceptance | Synthetic 95% minimum; normal split 95%, downside 93.5%; other gates stipulated | Predeclared checks, held/failed/unknown cases and independent reviewer |
| Included economics | Application and seats separate; USD6,800 transition | Disjoint rows, amounts/units, price/usage evidence and missing charges |
| Avoidable outflow | USD200/month assumed normal difference; six-month total increases USD5,600 | Cancellable amounts, commitments, timing and no-credit result |
| Transition/recovery | Not executed; retained dependencies remain | Trial scope, stop path, in-flight recovery and operator handover |
| Contract retirement | None authorized | Exact contract/objects, continuing consumers, effective date and approval/readback |
| Next decision | Obtain bounded task/access/evaluation evidence; do not cancel seats | Retain/trial/split/migrate/hold, owner and evidence needed |
11. Finish only when the remaining work is visible
The comparison is ready for its named decision when task and contract inventories reconcile; required route/feature facts are current; processing/access prerequisites are resolved for the proposed action; acceptance and operating evidence are independently reviewable; costs include retained and transition obligations; and removed spend has a supported effective date. A decision to retain or hold is a completed comparison outcome when its reasons and next action are recorded.
For this fictional packet, the result is to retain employee products and request a smaller application trial. The six-month cash comparison does not justify the proposed split on cost alone. A separate benefit might justify it, but that benefit needs its own evidence rather than an invented dollar value. No actual contract removal or migration is approved by the arithmetic.
Choose one capability and ask finance for its current product/API contract rows. Have engineering identify every state/tool dependency and the exact proposed route. Fill the record before commissioning paid evaluation. If those owners need implementation help, Ampity can scope an AI integration engagement around the missing evidence and agreed deliverables. Reading this resource or using its worksheets should not require contact information.
Related services
AI Product Integration & OpenAI Development Services
Embed AI capabilities into your existing products without rebuilding them. Integration architecture, latency strategy, fallback design, cost controls, and operational tooling from day one.
AWS Consulting Services for Architecture, Cost & Migration
AWS consulting for architecture and Well-Architected reviews, cost optimisation, FinOps and migration. Plan changes around workload evidence and recovery needs.