What Should an AI Workflow Do During a Provider Outage?
Choose safe delay, tested fallback or a bounded manual path during an AI outage. Preserve task state, reconcile uncertain writes and control backlog recovery.
Preserve the task before trying to preserve the response time
During an AI provider outage, choose the recovery path from the business task's current state. A read-only question may use a tested fallback. A draft may wait in a bounded queue. An action with an uncertain external effect needs reconciliation before another attempt. One global “switch to backup” rule cannot safely cover all three.
The goal is not to make every request appear successful. It is to retain a truthful record of what happened, prevent repeated effects and give the requester an understandable next step. A delayed answer with a clear status can be better than an immediate answer that silently changed the task's meaning or authority.
This article proposes an outage runbook for engineering teams. The worked example is synthetic, not an Ampity customer incident. Choose thresholds from your own latency requirements, observed failure patterns and operating capacity. No retry count or queue size in a generic article should become your production policy without testing.
Start by defining which parts of the service can continue without generation. Existing records, deterministic validation, search navigation and task-status lookup may remain useful. Keep those paths independent where practical. An unavailable model should not prevent a user from checking whether an earlier business action actually completed.
Distinguish dependency failure from task failure
A model endpoint timing out establishes that the caller did not receive a usable result within its deadline. It does not establish that every tool action associated with the task failed. Keep transport state, generation state and business-effect state separate in the task record.
For example, a scheduling assistant might have submitted an update before losing its final model response. The interface cannot infer that no schedule changed merely because the assistant stopped speaking. Show the update as pending verification, consult the scheduling system and keep the original operation identity through recovery.
Classify failures before choosing a response. Authentication errors, malformed inputs, quota exhaustion and a sustained dependency outage require different interventions. Repeatedly sending an invalid request is not recovery. A backup model will not repair a broken application credential or a missing source document.
Use an application-owned task identifier and separate attempt identifiers. Record the stage reached, any downstream receipt, the selected configuration and the recovery decision. Keep sensitive payloads out of general logs. Operators need enough evidence to reconcile state without giving everyone access to the source data.
Select a degraded mode for each task family
Write a small decision table before the outage. The useful distinction is what can be preserved under the fallback, not whether another provider happens to answer. Assign an owner to every path, including a deliberate decision to stop accepting new work.
| Task state | Candidate outage behavior | Condition that must hold | | --- | --- | --- | | Read-only answer, no effect requested | Tested fallback or unavailable message | Evidence, permissions and answer contract remain valid | | Draft generation not yet started | Queue or explicit manual preparation | Deadline, queue capacity and data handling permit delay | | Proposed change awaiting approval | Hold the proposal | Do not expand or reuse authority after inputs change | | External write sent, outcome unknown | Reconcile with the receiving system | No second write until recovery policy permits it | | Time-sensitive task with no safe substitute | Reject or escalate through a defined channel | Explain the unmet deadline rather than claiming completion |
Manual preparation is not an imaginary support team. If nobody is assigned and available, do not offer immediate human handling. State the actual response window or let the user leave the task for later. The runbook must reflect the staffing you have, not the staffing the interface implies.
Fallback also needs a narrower scope when appropriate. A substitute accepted for summarizing public material may be unsuitable for interpreting restricted documents or proposing writes. Disable unsupported task families independently. Maintain one clear status vocabulary across the original and degraded paths so that queued, prepared and completed do not become interchangeable.
Bound retries and isolate the failing dependency
Microsoft's Retry pattern distinguishes transient failures from faults that should not be retried and highlights idempotency and layered-retry risks. Apply that guidance to the specific operation. A model request, a retrieval call and a business write do not share the same retry safety.
Set an end-to-end deadline as well as per-attempt limits. Count retries performed by SDKs, gateways and workers, not just the loop visible in application code. Otherwise a seemingly small retry policy can multiply into a long wait and a large volume of failed calls.
The Circuit Breaker pattern describes temporarily blocking calls after failures and allowing limited probes during recovery. That can protect capacity. It does not determine whether an AI fallback is semantically acceptable or whether a previously submitted business action may be repeated.
Scope failure isolation carefully. An issue affecting one model route or task configuration should not necessarily block unrelated status lookups or healthy task families. Conversely, a shared credential failure should not be hidden by treating every request as an independent transient fault. Test the classification and the dependency boundaries together.
Worked example: an intake queue during a sustained outage
Consider a hypothetical document-intake service receiving 60 items an hour. Its approved degraded mode stores submissions for later preparation, with no automatic posting to a financial system. An outage lasts two hours, creating 120 queued items. These numbers illustrate planning arithmetic, not a recommended queue limit.
After recovery, suppose verified processing capacity is 90 items an hour while new arrivals remain at 60. Net drain capacity is 30 an hour. Clearing the 120-item backlog takes about four hours if those rates stay stable. “The provider is back” therefore does not mean the service has caught up.
If the oldest items have a deadline earlier than that, the operator needs a specific response: temporary intake limits, an available review route or an honest missed-deadline notification. Adding workers without checking provider quota, downstream capacity and review demand can move the bottleneck rather than remove it.
Before replay, recheck whether queued items remain valid. A document may have been replaced, the requester may have cancelled or the relevant approval may have expired. Retain the submission history, but do not treat an old queue entry as permanent authority. Reconcile any uncertain downstream effects separately from never-started generation work.
Tell the user what is delayed and what is known
An outage message should identify the affected capability, the task's status and the available choice. For a queued draft, say that preparation is delayed and no business update has been confirmed. For an uncertain update, say that the system is checking the receiving record. Those are different situations and deserve different messages.
Do not promise an automatic email or a recovery time unless the delivery mechanism and operating policy support it. A task-status page can be useful without collecting contact information. If notification is optional, make it optional and explain its purpose separately from access to the service.
Give cancellation a defined meaning. Removing never-started work from the queue is different from cancelling an external operation already accepted by another system. The cancellation result must describe what was stopped and what still requires verification. Do not clear the evidence simply to make the interface look clean.
Keep repeated submissions visible to the recovery logic. A frustrated user may refresh or submit again. Where the business operation supports it, bind attempts to a stable operation identity and show the existing pending task. A new browser request should not silently create a second external action for the same unresolved intent.
Limitations: a fallback is not a universal availability guarantee
Two providers may depend on shared networks, application code, credentials or data pipelines. A second endpoint does not prove independent failure modes. Capacity, approved data handling and behavior compatibility need their own checks. A fallback that cannot meet those requirements should remain disabled for the affected task.
Queuing is also a promise with a cost. Define maximum age, admission limits, cancellation handling and disposal of work that can no longer be performed. An unbounded queue can make the interface look available while accumulating tasks the organization cannot complete responsibly.
Recovery exercises cannot predict every outage. They can expose preventable errors: retry amplification, abandoned task records, unsupported fallback semantics and uncontrolled backlog replay. Keep incident review focused on which assumption failed and which acceptance test now needs to change, rather than declaring the design resilient because one rehearsal passed.
Rehearse the transition back, not just the outage
Choose one task family and run a controlled exercise without live business effects. Interrupt generation before any tool call, after a proposed action and after a simulated external receipt. Verify that each interruption leads to the correct state and that uncertain effects do not trigger blind retries.
Then restore the dependency gradually. Check that probes are bounded, queued work respects current permissions and deadlines, and fresh traffic does not starve older valid tasks. Record backlog age, completion outcomes and reconciliation failures alongside dependency error rate. A healthy endpoint is only one recovery signal.
Use fallback output contracts to assess substitution and tool timeouts and duplicate actions to review uncertain writes. The action-recovery playbook provides a structured rehearsal. For implementation help, explore production AI systems or share an outage-design question. Reading these resources does not require sharing personal details.