Split an AI Migration by Capability, Not by Endpoint
Keep provider-bound work on its existing route while evaluating a bounded Bedrock capability. Record state ownership, held work, routing evidence and the full...
An AI product can keep provider-bound capabilities on its current integration while evaluating a narrower capability on Bedrock. Make the split at a complete application task: its required inputs, state, accepted output, processing permissions and failure behavior. A shared API shape is insufficient. If the candidate needs state that remains available only through the direct provider, hold that capability or retain its existing route.
Consider a fictional support product with three operations. It summarizes a supplied passage, answers questions through hosted file search, and continues existing provider-held conversations. The team proposes moving only new passage summaries. Hosted retrieval and existing conversations remain on the source. That is a deliberate partial migration with two operating dependencies, not completed provider exit or automatic outage protection.
This article gives the application owner a capability-routing record and a worked split-cost example. All task aliases, evidence records, prices and outcomes are synthetic. No provider request, model evaluation, account read, export or deletion was executed. A local checker tests supplied records only. It cannot establish that a real account may process the data or that either model will produce an acceptable answer.
1. Pin a capability to an exact documentary route
Route A in this example is direct OpenAI Responses at https://api.openai.com/v1/responses, using gpt-4.1-2025-04-14. The model reference identifies that snapshot and Responses support. The GPT-4.1 feature section lists file search. These observations support the comparison scope, not present account access or a promise of indefinite availability.
Route B is Bedrock Runtime Converse at https://bedrock-runtime.us-east-1.amazonaws.com, using the US profile us.anthropic.claude-haiku-4-5-20251001-v1:0. The Haiku 4.5 card lists this Runtime API/profile combination. From us-east-1, its listed US destinations include us-east-1, us-east-2 and us-west-2. The source endpoint does not establish single-Region processing. Account prerequisites, actual permissions and acceptable locations remain unobserved.
The same card does not list Responses support for this model. Bedrock has a separate Responses API for supported model/endpoint combinations, including stored response state. Do not generalize the chosen Converse boundary into “Bedrock has no state,” or combine one model's API column with another model's route. AWS's compatibility matrix is a discovery aid; use the endpoint-specific card before choosing a candidate.
Pin the request and application revisions with the route. This specimen is synchronous, text-only and has no generated tool proposals or external business writes. It excludes images, streams, background work, prompt-resource indirection and model-directed routing. Adding any of those is a changed contract requiring evidence, even if the model identifier stays the same. Exact feature support and observed application behavior are separate review questions.
2. Separate the three operations before changing traffic
The summary capability accepts application-owned text from revision D29-r1. Its task rule asks for a short draft that preserves an explicit exception and identifies D29-r1. If the source says a support plan covers weekdays except public holidays, a summary that omits the exception fails. Neither fluency nor a successful provider ending establishes acceptance. The model output remains a draft for the product's existing review boundary.
The knowledge capability asks a question whose evidence must be retrieved from the existing hosted store. OpenAI's file-search guide describes uploaded files, vector stores and a hosted tool. In this proposed split, that dependency stays with route A. It is not implemented on B, which is narrower than claiming Bedrock has no retrieval services. A later retrieval redesign would need a new capability contract and a comparison of evidence coverage.
The continuation capability depends on an existing provider conversation. OpenAI's conversation-state guide describes durable conversation objects and their items. Route B's Converse request accepts messages and a model/profile identifier; it does not import a direct provider conversation by accepting its identifier. The state inventory and reconstruction decision belongs to the stored-state companion, not a silent router conversion.
- New passage summary, summary-v1
- Application context and revision are available without provider-held objects. B is the proposed candidate, conditional on separate task, permission and operating evidence. No source conversation, File or vector-store reference crosses this route.
- Hosted knowledge answer, knowledge-v1
- The source project's vector-store dependency remains required. Retain A under its existing independently controlled operation. If that dependency is unavailable, hold the answer or offer the underlying permitted documents, not an evidence-free B response.
- Existing conversation, continuation-v1
- Preserve source affinity and its namespaced conversation reference. Retain A. Reconstructing a new task from permitted application records would be a separately reviewed change, not continuation merely because the user sees one chat window.
The split is useful only if the summary can stand on its own. If every “summary” secretly performs hosted search first, its real dependency still includes A. If a customer request needs both retrieval and summarization, record those suboperations and their completion rules. Do not advertise the whole request as migrated because one inference step moved. The example intentionally avoids that composite workflow so its boundary is inspectable.
3. Route from trusted application state
Select the capability from the authenticated application's operation, not from a label supplied by a model or arbitrary caller. Bind the task, capability revision, state revision, route tuple and reviewed contract reference before dispatch. Keep the capability decision observable in the task record. An operator should be able to explain why a hosted-search request remained on A without reading a prompt or guessing from a billing line.
The diagram shows the proposed split. Solid arrows indicate request selection, not copied state or successful responses. Source-held objects stay inside the retained route's responsibility. The application owns task identity, the accepted source revision and output acceptance. It also owns held-work status when a route cannot meet its contract. No arrow between routes means no implicit conversation transfer and no automatic fallback.
Proposed capability routing, with no executed requests. A and B are bound to the exact provider/API records above. Separate desktop and mobile compositions show the same capability boundaries.
Version the routing policy separately from a model prompt. A rollback changes admission for new eligible work under a reviewed policy; it must not rewrite the stored route of an in-flight task. Retain old policy and attempt records. If an old attempt returns after a route change, its result still belongs to that attempt and its original acceptance rule. The continuity paper owns the broader incident and exit policy.
The offline specimen accepts three exact profiles and only a not-dispatched attempt. Its favorable status is SUPPLIED_ROUTE_REVIEW_ONLY; execution authorization, provider validation and dispatch remain false. A source conversation attached to a summary causes HOLD instead of being discarded. Another capability's evidence reference, a changed Region/profile or a different task's state also holds. These are checks on supplied declarations, not proof of their origin or truth.
4. Make unavailability a visible disposition
Partial migration and runtime fallback need different evidence. The team's ability to operate summaries on A before migration does not prove that every B attempt may be replayed on A after an interruption. A timeout can leave remote work and accounting uncertain. For this no-write profile, it can still produce duplicate charges or a late answer incorrectly attached to a new attempt. Keep the uncertain attempt visible and follow the separately reviewed recovery policy.
The specimen holds every already-started or uncertain attempt. It does not infer a pre-dispatch failure from an HTTP category or exception name. A future production implementation may support narrower retry or fallback cases, but it must establish attempt status, current data permission, remaining deadline and budget, a fresh attempt identity and eligible output handling. A second route's health alone provides none of that evidence.
Three useful adverse cases change the proposed disposition. First, a knowledge request is mislabeled for B while retaining a vector-store reference: HOLD the route mismatch. Second, a summary contains only a source conversation ID: HOLD the missing application context boundary, even though the string is nonempty. Third, a B summary has an uncertain prior attempt and A is healthy: HOLD automatic replay. The recovery owner can investigate; the checker never invokes either provider.
Capability-specific degradation can preserve useful service. A user may still open a permitted source document, save a draft or see that a conversation is awaiting the retained route. Those deterministic paths require their own current access checks. Do not present a generic generated answer as the original knowledge capability when its required source is unavailable. An honestly pending task is preferable to changing the acceptance rule during a disruption.
5. Account for the retained route and extra work
Use one declared population and window when comparing costs. The fictional C29 window contains 100 distinct tasks: 70 summaries, 20 hosted knowledge answers and 10 continuations. Its old-route inference-only baseline stipulates four US cents per summary, twelve per knowledge answer and eight per continuation. Thus 70 × 4 + 20 × 12 + 10 × 8 = 600 cents. These are teaching amounts, not provider prices, measured invoices or a total ownership baseline.
The partial split stipulates three cents per B summary attempt while retaining the other two categories' costs. Initial inference is 70 × 3 + 20 × 12 + 10 × 8 = 530 cents. Five extra summary attempts add fifteen cents. Additional dual-operation effort is allocated 150 cents for this same window, giving 530 + 15 + 150 = 695 cents. The included split amount is 95 cents above the inference-only baseline despite the lower assumed summary rate.
There are 105 attempts but still 100 logical tasks. Stipulate 90 accepted tasks and ten held tasks in this illustrative split window. Its included cost per accepted task is exactly 695 / 90 cents, approximately 7.7222 cents. The baseline has no supplied accepted-outcome count, so a comparative accepted-result ratio is UNKNOWN. Do not divide the old 600 cents by an invented 100 successes or count retry attempts as additional accepted work.
The companion preserves integer cents and the exact numerator/denominator. It returns no ratio when accepted count is zero; incurred costs do not disappear. It rejects absent amounts instead of treating them as zero. A real zero price is representable, but would need a documented basis outside the model. Its aggregate extra-attempt count cannot exceed the summary-task count. That is a local specimen limit, not per-task retry enforcement, recommended retry policy or a provider quota.
Actual costs need a different ledger. Bedrock pricing distinguishes model, operating mode and usage dimensions. OpenAI pricing separately describes model tokens, file-search storage and tool calls. Keep retained storage, failed attempts, duplicated context, logs, evaluation, integration work and operator review visible. Rate changes, caching assumptions or a different accepted population can change the result. Taxes, currency conversion, migration build cost and the common pre-existing operating baseline are excluded from C29. No savings or ROI conclusion follows.
6. Keep data and observability boundaries separate
Moving summaries does not authorize deleting source Files, conversations or indexes still needed by the other capabilities. It also does not settle the retained provider's account, access or retention obligations. The source-state companion owns object-specific disposition. Keep the dependency owner and continuing consumers in the split record so a cost-reduction action cannot accidentally remove the retained route's evidence.
“Application-context summary” describes the input dependency, not a no-retention guarantee. Bedrock model invocation logging is optional and can record inputs and outputs for the selected Runtime operations. Application traces, diagnostic exports and accepted history are additional copies. Inspect actual settings and permissions before a trial. This article neither enables logging nor establishes a whole-system deletion or residency result.
Record task and attempt aliases, selected route revision, relevant evidence references, native request identity when actually obtained, disposition and route-specific usage. Avoid placing sensitive passages or credentials into broadly accessible routing logs. A model's text is not a trusted source for its own billing, permission or destination record. Unknown usage remains unknown until the appropriate receipt or account observation resolves it.
Useful checks include source-bound requests unexpectedly sent to B, tasks with no matching state revision, results accepted under another attempt's policy, and cost rows outside the comparison window. A route-level success rate cannot establish preservation of the weekday exception. Conversely, correct supplied answer text cannot establish that a deployment used the intended Region/profile. Keep answer, routing, permission and cost evidence separately attributable.
7. Fill one capability record before expanding the split
The filled record describes the proposed summary capability. Its roles and evidence references are fictional. The blank record uses the same labels so an application owner can replace each assumption with a permitted evidence reference, without copying private data into the worksheet.
- Capability and task
- summary-v1, fictional T29-S; new passage-summary draft preserving the weekday exception and D29-r1 reference. No external business write.
- State ownership and dependency
- Application-context revision D29-r1, linked to T29-S. No native provider references. Hosted knowledge and continuation remain distinct A capabilities.
- Exact selected route
- B: the Runtime Converse endpoint, Haiku 4.5 US profile and us-east-1 source specified above. Fictional account alias; actual access and permitted destinations UNKNOWN.
- Policy and evidence binding
- split-r1 plus fixture-contract-r1 for summary-v1/B. Supplied, not observed evaluation or approval. Real context completeness and output acceptance remain untested.
- Attempt and recovery boundary
- Not dispatched. Unknown or started attempts HOLD in the specimen. No automatic A fallback; separate owner review for any actual recovery.
- Cost population and exclusions
- C29 teaching window includes retained A work, five extra B attempts and allocated dual-operation effort. No quoted rates, total ownership claim or measured saving.
- Stop and accountable next action
- Integration owner requests independent contract review. Stop on missing state, wrong route, invalid permissions or changed acceptance. Any paid trial needs separate scope, budget and authorization.
- Capability and task
- Name the operation/revision, task identity, accepted output and excluded effects. Explain whether it is a complete task or one suboperation.
- State ownership and dependency
- Identify authoritative input revision, application/provider ownership, namespaced native dependencies and retained consumers. Mark incomplete evidence.
- Exact selected route
- Record provider, endpoint/API, model/version/profile, source Region, account/project scope and applicable feature/placement evidence.
- Policy and evidence binding
- Bind the routing revision, capability, adapter, acceptance rule and evaluated evidence. Separate documentation, supplied assumptions and observed results.
- Attempt and recovery boundary
- State what has been dispatched, returned, accepted or remains uncertain. Name held-work behavior, separately permitted retries and rollback limits.
- Cost population and exclusions
- Use one window/population; retain both providers, additional attempts, operational costs and accepted-result denominator. Identify missing amounts and excluded costs.
- Stop and accountable next action
- Name owners of missing evidence, independent review, stop conditions and changed bindings that reopen the decision. Record actual execution authority separately.
8. Challenge the split before treating it as an operating option
The offline companion contains the inspectable classifier, synthetic packet, adverse tests and cost arithmetic. It has no provider SDK, request builder, answer grader or deployment integration. A fabricated internally consistent packet can receive its review-only result. The supplied capability is assumed to come from trusted application state; this module cannot authenticate that assumption or inspect content behind a revision alias.
Retain the direct integration if the eligible capability is too small to justify two operating paths. Rebuild retrieval only when that separate decision is justified. Evaluate a different exact API/model combination if it meets the whole capability more naturally, but do not borrow old evidence for the changed path. Existing evaluation-expiry guidance explains when a prior result stops supporting that proposed decision. Tool, streaming and image contracts remain outside this text-only example.
Choose one real capability and complete the blank record. Have a second engineer challenge a retained-state request mislabeled as a summary, a wrong route binding and an uncertain prior attempt. The application owner should then decide whether the remaining gap is data reconstruction, task evaluation, operating permission or simply an uneconomic split. An AI integration review can turn those specific gaps into a bounded implementation scope. Preserve the current supported path while the candidate's evidence remains incomplete.
Related services
AI Product Integration & OpenAI Development Services
Embed AI capabilities into your existing products without rebuilding them. Integration architecture, latency strategy, fallback design, cost controls, and operational tooling from day one.
AI Observability, LLM Monitoring & Governance
LLMOps consulting for AI observability, LLM monitoring, evaluation and guardrails. Review production answer quality, operating failures and cost evidence.