AI Provider Continuity and Exit: Keep the Business Task in Control

Decide which AI tasks can use provider fallback, which must wait and how to preserve business state during a planned exit. Includes compatibility evidence, capacity...

audience="Engineering leaders and application owners whose business workflows depend on externally hosted AI." decision="Which tasks may continue during a provider failure, and what must remain under application ownership to make an eventual exit practical." position="Pre-approve substitutes for specific task contracts. Preserve state and authority outside the provider, and keep uncertain effects stopped until reconciled." scope="A proposed engineering decision framework with hypothetical examples. Not a customer implementation, availability guarantee, procurement assessment or legal opinion." outputs={[ 'A task-family continuity register', 'An evidence-backed compatibility ledger', 'A degraded-service and capacity policy', 'An application-owned exit inventory', 'A staged migration and reconciliation plan', 'A review checklist for switching and returning', ]} />

Executive summary

An application can reach a second model and still be unable to complete its business task safely. The replacement may interpret an exception differently, omit evidence, request a different tool or process data in an unapproved location. A successful response proves that an endpoint answered. It does not prove that the replacement satisfies the operating contract. Provider continuity therefore starts with the work that must remain possible, not with the number of provider logos in an architecture diagram.

This paper separates three decisions: retrying a transient failure, continuing a limited task through an evaluated substitute, and leaving a provider through a planned migration. They share some engineering components but have different evidence requirements. An incident is a poor time to discover that the alternative cannot interpret stored sessions. A planned exit is a poor reason to replay unresolved commands against a new model. Application-owned state lets the team make these decisions without asking a provider conversation to serve as the authoritative business record.

The recommended approach is selective continuity. Define task families, identify the effects they can produce and evaluate candidates against the same acceptance rules. Preserve a useful non-generative path where possible. Keep unsupported tasks visibly pending rather than fabricating completion. An exit plan then covers more than API adaptation: source data, permissions, retrieval indexes, sessions, evaluations, capacity, retained records and evidence of final effects. The result should be a defensible operating choice, not a promise that every model is interchangeable.

Start with the customer obligation

Ask what the customer needs during a disruption. A person reviewing an application may still need access to uploaded documents and the status of an existing request, even if a generated summary is unavailable. Another person may need a confirmation that a payment was recorded. These obligations are different. Document access can sometimes continue through deterministic software. Payment confirmation needs authoritative evidence from the payment system; a replacement model cannot establish it by making a plausible statement.

Write the minimum useful service for each family before designing fallback. Include the deadline, allowed degradation, required evidence, permitted data and owner of a final decision. A shortened answer may be acceptable for navigation but unacceptable for explaining a contractual exception. A delayed classification may be tolerable until a review window closes. An expired business deadline may make a queued task invalid even after technical capacity returns. The register must capture these conditions rather than treating every pending prompt as useful backlog.

Measure disruption in customer terms: valid requests not accepted, requests awaiting review, information unavailable and effects whose status is unknown. Model latency and error rate remain necessary diagnostic signals, but they do not fully describe the obligation. Decide who may change degraded-service policy and how users learn its limits. A continuity plan is useful only if someone can explain what remains available, what is held and what evidence will allow normal service to resume.

Separate retry, fallback and exit

Retry repeats an eligible attempt under a bounded policy. It addresses a failure expected to be transient. Fallback changes an execution path, potentially changing semantics, data handling and capacity. Exit changes the longer-term dependency and may require rebuilding derived state. Calling all three failover hides their different risks. A common client library can coordinate requests, but it cannot make these decisions equivalent or remove their effects on the business contract.

The Microsoft Retry pattern describes bounded strategies and warns that repeated operations can duplicate effects when an earlier result is lost. Apply that principle across providers as well as within one provider. Switching the generator does not undo a command already sent to a business system. A retry budget should cover application, gateway and SDK layers together; otherwise apparently small limits multiply into excessive attempts and delays.

Exit normally allows time for inventory, tests and staged exposure. Runtime fallback must already have that evidence for its narrower scope. A candidate that can replace a provider after several weeks of migration is not necessarily available as today's outage substitute. Record readiness separately: designed, tested in isolation, evaluated for a named family, provisioned for its expected load and authorized for use. These states prevent a procurement alternative from being presented as an operational safety net before the application can actually use it.

Classify failures before changing the route

Some failures indicate temporary congestion. Others indicate invalid requests, missing permissions, revoked credentials or exhausted spending allowances. Sending the same task elsewhere can conceal the cause rather than restore the intended service. For example, an authorization failure should not trigger a provider route that bypasses the organization's restriction. A malformed schema should lead to a configuration investigation, not an unvalidated free-text response that the application accepts as equivalent output.

The Claude error documentation distinguishes request, permission, rate-limit, timeout and overload errors, and explains that streaming failures can occur after an initial successful HTTP response. This is a concrete reminder to inspect the completed response, not just connection establishment. Capture bounded diagnostic information and request identifiers. Do not place complete sensitive prompts into unrestricted incident records merely because engineers need to investigate a route change.

Define actions by failure class and task class together. A known pre-execution refusal may allow an authorized new attempt. An interrupted stream requires an explicit incomplete-result state. A tool call with an unknown effect requires reconciliation. A configuration failure needs its owner. Avoid a universal timeout threshold that silently changes providers for everything. Continuity depends on knowing which part failed and whether the task's evidence, permissions and effect history are still usable.

Make eligibility a business decision

Before admitting a task to fallback, answer three questions. Is the prior effect state known? Is the candidate approved for this exact task family and data scope? Can the application still enforce the output and action contract? A yes to only the last question is insufficient. Validation can reject bad output but cannot authorize new processing locations or prove that an earlier command did not execute. These checks belong to admission, before transferring the task.

The proposed decision tree is deliberately not a routing topology. It shows why a task stays held even when a second endpoint is reachable. Unknown effects remain outside fallback admission. Eligible tasks still require current evidence, permissions, bounded load and output checks. The diagram does not imply that every read-only task is harmless: disclosing the wrong information is itself an important consequence. Nor does an evaluated write-capable task skip the application-owned command boundary.

Store eligibility as a versioned policy with a reason for each exclusion. Operators should not have to infer it from a prompt. The allowed set can be narrower than ordinary service: for example, source-linked summaries may continue while automated dispositions are held. Tell users about that distinction. A visible limitation is usually more defensible than a seamless-looking interface that changes the meaning or confidence of its answers without notice.

Build a compatibility ledger, not a feature checklist

A feature comparison asks whether two services support tools or structured output. A compatibility ledger asks whether the application can obtain an accepted result from each candidate under its actual conditions. Include required fields, exceptional results, source references, tool argument types, permission boundaries, context limits and processing restrictions. Record the tested adapter and configuration versions. Similar feature names are discovery signals, not evidence of equivalent behavior or readiness for a particular task.

Use a compact record per family: business acceptance rule, candidate configuration, independent fixtures, observed failures, capacity limit, permitted data, unresolved dependencies and the owner approving exposure. Retain examples that distinguish candidates rather than only ordinary examples on which they agree. Include missing evidence, contradictory records, ambiguous names, revoked access and attempts to change scope. A candidate may satisfy one family while remaining excluded from another. The ledger should preserve that uneven result rather than averaging it away.

Review compatibility when the application contract or relevant provider behavior changes. A prior pass does not cover a new tool, schema, retrieval path or data class. Link the ledger to the AI change release-governance paper so routing changes receive the same scrutiny as other behavior changes. Keep evidence that a candidate failed as well as evidence that it passed. Without negative results, the document becomes a sales comparison rather than a useful operating constraint.

Keep structure, meaning and authority separate

Structured output helps software parse an answer. It is not a substitute for checking the answer's business meaning. A valid object can contain a wrong recipient, an unsupported conclusion or a stale reference. Normalize candidate responses into an application contract and validate required evidence there. Do not normalize away significant exceptions. A refusal, incomplete output or unsupported citation should remain visible, not become an empty success object that downstream code treats as completed work.

The Claude structured-output guidance documents schema limitations and cases where refusal or token limits can prevent the expected output. That illustrates why adapters need explicit exceptional states. The application should reject or hold unsuitable results independently of any provider's success flag. A different candidate requires its own current contract review; this documentation does not establish another vendor's guarantees or prove that providers interpret a shared schema identically.

Authority belongs to deterministic application controls. A model may propose a command, but the application verifies the actor, target, current permissions, record version and any required approval. Keep those checks unchanged when switching generators. If a replacement requires broader tool access to succeed, that is a new design and authorization decision, not a continuity improvement. Evaluate semantic quality and authorization failures separately. A favorable answer-quality average cannot compensate for even one prohibited effect.

Evaluate against consequences, not answer resemblance

Two answers can look different and both be acceptable. They can also look similar while one contains a decisive error. Evaluate the consequences defined by the task contract: source applicability, field correctness, allowed uncertainty, correct disposition and valid proposed command. Where judgment is required, use qualified reviewers with a written rubric. Keep representative task families and adverse cases visible. A single overall score can hide a candidate's failure on the very exceptions that make continuity important.

For a hypothetical document-summary family, require a link to each decisive source, separation of known facts from absent evidence and no automatic record changes. For a hypothetical disposition family, add explicit rule checks and a review boundary. A candidate passing the summary family is not thereby approved for disposition. These examples are proposed evaluation contracts, not measured results. They show how scope can expand through evidence rather than through the assumption that a generally capable model handles every adjacent task.

Test operational behavior too: interrupted streams, output limits, late arrivals, timeouts, denied tools and responses after cancellation. Validate that rejected output cannot reach a command path. Compare outcomes on equivalent task populations and record reviewer effort, delay and unresolved exposure. This keeps the evaluation connected to the workflow unit-economics framework. A substitute with lower provider charges but much greater checking effort may offer useful emergency capacity without being an attractive permanent replacement.

Design degradation before buying redundancy

Useful degraded service may require no replacement model. Preserve document viewing, deterministic search, request status, previously verified records and the ability to save a draft where those functions remain safe. Remove unavailable generation controls or describe their limitation clearly. Do not let an interface suggest that a delayed task has completed. The best alternative depends on the obligation: a stable status page can be more valuable than a new model producing an uncertain explanation of that status.

Evaluate three operating choices explicitly. A single provider with a well-designed held-work path has less adapter complexity but limited generation continuity. A pre-evaluated substitute for narrow families adds testing and capacity commitments while preserving some work. A broader multi-provider implementation can offer more options but carries additional semantic, data and operational dependencies. None is universally best. Choose based on disruption consequences, manageable scope, team capability and the evidence required to keep the service trustworthy.

A self-hosted model is another possible design, not a free emergency button. It requires model suitability, compute, security, deployment, monitoring and incident capability. Do not assume it will be economical or instantly available. An incident plan should use an already tested operating path, not a hypothetical future build. Make the unavailable functions and recovery owner explicit so the customer receives an honest experience and the operations team is not forced to invent policy under pressure.

Check independence and data boundaries

Two routes may share the same model provider, cloud, gateway, identity service, network path or account allowance. Their visible endpoints can differ while a common dependency still prevents work. Draw a dependency inventory and identify the failure each alternative is intended to address. Capacity relief within one provider can be worthwhile, but it is not the same as independence from that provider. Assess common dependencies before assigning an alternative an availability objective.

The Amazon Bedrock cross-region guidance describes inference profiles that route a model across destination regions, including geographic and global options. This is a capacity-routing mechanism, not proof of a different provider or an application recovery contract. Inspect the actual profile, allowed destinations and organizational policy. A geographic label alone does not establish every contractual obligation, and a globally available route may be unsuitable for a task whose processing location is restricted.

Inventory data handling for each candidate and feature: inputs, uploaded files, sessions, outputs, logs, caches and retained diagnostic information. Verify account-specific terms and configuration with the appropriate owners. Approval for one service is not blanket permission for another. Keep processing restrictions in admission policy rather than a comment in a document. This paper supplies engineering questions, not legal compliance advice; specialized contractual, privacy and regulatory conclusions need qualified review for the actual application and jurisdictions.

Budget substitute capacity and held work

An available candidate can still be unable to carry the displaced workload. Distinguish nominal account access from provisioned or observed capacity for the intended task mix. Context length, tool interaction, output size and latency can change resource demand. Give the substitute its own admission limit and retry budget. Protect unaffected deterministic operations from the queue and resource usage of degraded generation. A continuity feature should not turn a provider problem into a wider application outage.

The Microsoft Circuit Breaker pattern describes blocking unsuccessful calls and admitting limited recovery probes rather than flooding a recovering dependency. Use that principle together with task-specific admission. The provider's apparent recovery is not sufficient to release every held task at once. Keep the operating decision separate from the mechanism that opens or closes a circuit. Probes establish limited evidence about the tested path, not all task families or business effects.

Consider a synthetic example with forty incoming eligible tasks per hour and substitute capacity for twenty-five. The queue grows by fifteen tasks per hour, before considering rework. After two hours, thirty additional tasks await service. If later usable capacity reaches fifty tasks per hour while arrivals remain forty, the net drain is ten per hour: three hours to clear that added backlog, assuming the tasks remain valid and no other work competes. This is illustrative arithmetic, not a vendor throughput claim. Use observed task mix and deadlines before making a service promise.

Preserve the task when an effect is unknown

An interrupted model conversation can be restarted from an application-owned record. An uncertain business action cannot simply be restarted as though nothing happened. Keep effect identifiers, intended payload, authorization version and authoritative status outside the provider session. If an outbound notification or record update might have succeeded, reconcile it with the relevant system before reissuing. A change of model does not change the identity of the business operation or make duplicate effects acceptable.

Preserve the distinction between a task, a model attempt and each effect. A task may involve several attempts and more than one command. One completed command does not establish completion of all remaining effects. Use stable operation identities where supported and reject reuse with changed intent. Where the destination cannot establish idempotent or queryable behavior, hold uncertain work for controlled investigation rather than treating a new provider route as evidence that the first action failed.

The action-recovery paper develops those contracts in detail. Continuity policy depends on them but has a different purpose: deciding which work can safely continue through another execution path. Keep the recovery record accessible even if provider tools or stored sessions are unavailable. If the application cannot determine the prior effect state without that provider, mark the dependency explicitly. It is an exit and continuity gap, not a reason to manufacture a confident recovery status.

Own portable state without promising universal portability

Application-owned state should contain the business request, task status, authoritative source references, actor and permission scope, contract version, accepted outputs and effect receipts. Preserve necessary evidence under a defined retention and access policy. Provider session identifiers, native message blocks and cached state can remain in an adapter-specific record. They may help diagnostics or efficient interaction, but they should not be the only means of reconstructing what the application was asked to do or what it changed.

This ownership view is not a claim that every provider session can be exported or replayed elsewhere. Some native artifacts may be unavailable, restricted or meaningless to another model. Build an explicit reconstruction from permitted authoritative inputs when appropriate, and evaluate it. Record what will not transfer, what must be rebuilt and what can remain archived. Do not copy hidden reasoning or proprietary internal state and assume it is a reliable or permitted substitute for a business explanation.

The abstraction should be as small as the application needs. Preserving a source reference and acceptance rule is often more useful than inventing a universal conversation format. Keep raw adapter details available to authorized investigators without exposing them as the domain contract. Application ownership reduces avoidable coupling; it does not remove legitimate model differences. A portable record creates options, while independent evaluation determines whether any particular option is acceptable for the intended work.

Treat retrieval migration as a separate workstream

Replacing the generator is different from replacing the embedding model, retrieval service or provider-managed document store. Identify each dependency independently. If the current retrieval path remains available and permitted, a generator substitute might use the same approved evidence. If retrieval is also unavailable, the substitute needs an evaluated alternative or a held-task path. A confident answer without the task's required evidence is not continuity merely because it arrives quickly.

The Amazon Titan embedding documentation distinguishes generating vectors from similarity computation performed by a vector database and describes supported output dimensions. That distinction matters for an exit inventory. Matching vector dimensions is not evidence that two models share a compatible semantic space. A planned embedding migration should rebuild and evaluate the relevant index using the candidate representation rather than silently comparing new queries against unverified old vectors.

Retain permitted source material, document versions, chunk boundaries, access metadata and deletion or revocation state needed for the rebuild. Re-evaluate ranking, evidence coverage and answer support. Check permissions at serving time as well as during ingestion. An old index snapshot may contain access that has since been withdrawn. Use bounded dual-index comparisons where appropriate, without treating duplicated storage as indefinite permission to retain data. The migration is complete only when the new retrieval path meets its contract and old copies have a documented disposition.

Plan exit as a controlled migration

Begin with a dependency inventory: generation, embedding, files, sessions, tools, batch jobs, fine-tuned artifacts if applicable, billing records and diagnostic retention. Assign an owner and exit treatment to each item. The treatment may be export, rebuild, archive, expire or delete subject to applicable obligations. Unknown items need investigation before an irreversible cutover. A repository search for provider imports is useful, but it will not reveal all operational state, account settings or workflow dependencies.

Define the cutover unit. It might be a new task family or a cohort of new requests, not every active conversation simultaneously. Preserve a configuration version on each task so in-flight work follows a known contract. Prepare a rollback path that can read the state created during migration, or state where rollback means holding work for reconciliation instead of automatic reversal. A migration that changes domain state beyond the old implementation's comprehension cannot be safely undone by changing an environment variable.

Run isolated evaluation, bounded shadow comparison where data handling permits it, limited accepted exposure and a wider release only after reviewing evidence. Shadow work must not issue real commands or send customer communications. Compare source coverage, exceptions, effect safety, elapsed time and operating effort. Keep an explicit stop condition and owner. There is no universal exposure percentage or duration; choose those from task consequence, detectability and recoverability rather than a copied rollout schedule.

Return deliberately and retire dependencies carefully

After an outage, choose whether to return, remain degraded or keep using an evaluated substitute. Do not make return automatic solely because a status page becomes green. Probe the actual application path and relevant task families with bounded load. Re-check credentials, quotas, schema behavior, required retrieval and permissions. A long disruption may have changed deadlines, source versions and approvals. Revalidate held tasks before admitting them, and keep uncertain effects on their reconciliation path.

For a permanent exit, reconcile remaining work before retiring credentials, webhooks or provider-managed resources. Establish which records must remain accessible for legitimate support or retention purposes and which copies should be removed. Keep evidence of requests and confirmed dispositions distinct. A deletion request is not proof that every retained copy is gone. Avoid deleting the only diagnostic record needed to resolve an uncertain effect, and avoid retaining unrestricted sensitive payloads merely for convenience.

Record the new steady-state owner, dependencies and acceptance evidence. Retire obsolete fallback rules so later incidents do not route traffic to a forgotten account. Remove secrets through the organization's controlled process, not by publishing configuration details in a migration report. Verify that monitoring and alerts now reflect the actual operating path. A completed exit should leave a smaller, understandable dependency set, not two partially active integrations whose permissions and obligations nobody owns.

Review the evidence before authorizing a switch

Use these questions as a review agenda, not as a compliance certificate. Each answer needs an owner and evidence appropriate to the task. An unanswered question can justify a narrower scope or a held-work policy rather than an immediate multi-provider build. The review should establish what the application can do today, what depends on a future migration and what it must not attempt under disruption.

  1. Which customer obligations remain available without generation, and which have a deadline that makes delay consequential?
  2. Which exact task families are approved for the candidate, with which adapter, output contract and evaluation evidence?
  3. Are processing locations, data features and permissions authorized for the task, including diagnostic and retained copies?
  4. Can the application determine prior effect status without access to the failing provider's session?
  5. What observed capacity, admission limit and combined retry budget protect the substitute and unaffected functions?
  6. Which retrieval, session or file dependencies require rebuilding rather than an endpoint change?
  7. Who can stop exposure, reconcile held work and authorize staged return or irreversible retirement?
  8. What evidence would show that the chosen continuity policy is failing customers despite acceptable endpoint metrics?

Document the decision and its exclusions in language that operations, engineering and the business owner can all use. Give it a review trigger: contract change, provider behavior change, new data class or failed rehearsal. The provider-outage playbook provides an execution drill for a policy already defined. The decision record should precede that drill. Otherwise the rehearsal only proves that traffic can move, while the harder question, whether the moved work remains acceptable, stays unanswered.

The practical next step

Choose one consequential task family and create its continuity record before expanding the provider set. Name the minimum useful degraded service, the effects that must remain held and the acceptance evidence required from a substitute. Inventory the state presently stored only with the provider. Test a candidate in isolation against adverse fixtures, and document both supported and excluded paths. This bounded exercise will reveal whether the immediate need is an adapter, better recovery evidence, a retrieval rebuild or simply a clearer customer experience during disruption.

Ampity's AI engineering services and agentic workflow work can support a scoped assessment of those boundaries. Define deliverables, dependencies and acceptance criteria for the actual system before committing to a build. The goal is not provider independence as a slogan. It is an application that can explain what it knows, protect what it may change and make a defensible continuity or exit decision when its AI dependency no longer behaves as expected.