Rehearse an AI Provider Outage

Run an isolated outage drill that tests bounded retries, truthful task status, approved fallback and controlled backlog recovery without repeating business actions.

trigger="A workflow depends on an AI provider, but the team has not demonstrated what users, queued tasks and uncertain actions experience when that dependency fails." owner="The application owner accountable for accepted business outcomes, with a named exercise controller." participants={['Application engineer', 'Platform or operations owner', 'Domain acceptance owner', 'Security reviewer', 'Independent observer']} prerequisites={['An isolated environment with synthetic inputs and effect sinks', 'Versioned task and fallback contracts', 'Visible task, attempt and effect identities', 'An approved exercise scope, abort control and recovery owner']} outputs={['A fault-and-expectation register', 'An observed outage and recovery timeline', 'Task and effect reconciliation evidence', 'An owned remediation list and an explicit release recommendation']} doneWhen={['Every fixture has an evidence-backed disposition', 'Retries and probes stay within declared limits', 'Unsupported fallback paths remain disabled', 'Recovery preserves permissions, validity and effect identity', 'No simulated uncertain write is blindly repeated']} />

Test the task through failure and recovery

Use this playbook to discover whether your AI-assisted application can fail honestly and recover safely. The useful result is not a screenshot of a backup model answering. It is a record showing how each business task was contained, communicated, reconciled and resumed, or deliberately left stopped.

The procedure is proposed engineering guidance, not a claim about a deployed Ampity customer system. The exercise examples are synthetic. Choose deadlines, failure thresholds, capacity limits and acceptance criteria with your application owner. A successful rehearsal supports a bounded recommendation; it does not prove uninterrupted availability or authorize a production experiment.

Start with one task family. Include a harmless read-only request, never-started queued work and a simulated action whose downstream outcome is uncertain. These states require different treatment. A fallback that preserves an answer contract cannot establish whether a business update already happened. Keep provider availability separate from task truth throughout the drill.

For the underlying choices, read AI provider outage degraded modes. This playbook turns those choices into testable operating steps. If you cannot isolate business effects or identify a person who can stop the exercise, complete the preparation first. Do not compensate for an unsafe harness by asking participants to be careful.

1. Authorize a narrow exercise and a real abort control

Owner: exercise controller with the application owner. Output: signed-off exercise scope. Identify the environment, task family, participants, start window and maximum duration. Specify which dependencies may be interrupted and which must remain healthy. List every excluded production credential, customer dataset, webhook and external write destination.

Use a stub or disposable effect sink for consequential actions. Demonstrate that the harness cannot reach the live system by checking network destinations, credential scope and a test receipt. A configuration label saying “staging” is not sufficient if its email sender, CRM connection or payment endpoint still points at production.

Define the abort command and who may use it. Aborting should stop injection and new test work, preserve task evidence and prevent uncontrolled replay. Rehearse that command before introducing faults. Tell observers where they can see its effect and which running operations may still need reconciliation.

Set stop conditions that do not depend on the controller's judgment in the moment: an unexpected outbound destination, missing task identities, growing unbounded retries, lost effect records or a breached isolation boundary. If any appears, end the exercise and retain the failure. Do not continue solely to complete the planned agenda.

2. Write expected behavior before injecting failure

Owner: domain acceptance owner with the application engineer. Output: task-state expectations. For each fixture, record the starting state, permitted degraded behavior, visible status and authoritative completion evidence. Decide what an acceptable delayed outcome means. An assistant's final sentence is not the source of truth for an external update.

Separate acceptance of intake from acceptance of completed work. A queue receipt means the system retained the request under a defined policy. It does not mean a draft was prepared, a review happened or a receiving record changed. The interface and operational dashboards should preserve those distinctions.

Write expectations for cancellation and expiry. A task cancelled before execution should not run during recovery. An approval tied to an older record revision should not authorize a changed action. An elapsed business deadline can make otherwise valid work useless. Include the status that explains each outcome to the requester.

Record the fallback boundary at this stage. If no substitute has been evaluated for the task's data handling, output structure and domain meaning, the expected path is hold or unavailability. Testing an outage is not an opportunity to introduce an unapproved provider. Keep the unknowns visible in the exercise register.

3. Build a fault register that distinguishes causes

Owner: application engineer. Output: reproducible fault fixtures. Include a connection failure, a delayed response beyond the application's deadline, a sustained overload response and a non-retryable request or credential error. Add a interrupted stream if the application streams results. Preserve the difference between a partial display and an accepted final output.

Claude's error documentation distinguishes authentication, permission, rate-limit, server and overload errors. It also describes errors after a streaming response begins. Use the documentation for the provider and access path you actually run. Do not reduce every failure to a generic server error or infer retryability from a status number alone.

Inject one fault at a time initially. Record its scope and duration so an observer can reproduce it. After individual cases pass, combine only those failures that answer a specific question, such as whether an SDK retry plus a worker retry exceeds the total task deadline. A chaotic test without a causal record is hard to diagnose.

Include the application's own broken credential or invalid input as a control case. The expected result is an owned configuration or validation failure, not an automatic traffic switch that hides the defect. Identify who receives the actionable signal. The runbook should distinguish a provider outage from a problem the application team must repair.

4. Count retries across the whole call path

Owner: platform engineer with the application engineer. Output: observed retry budget. Enumerate application, SDK, proxy, gateway and queue-worker retry policies. Record their deadlines, attempt limits and ownership. If a layer is outside your control, capture its observable behavior and the remaining uncertainty rather than treating its configuration as known.

Microsoft's Retry pattern emphasizes transient faults, idempotency and the risks of nested retries. In your exercise, measure attempts against one stable task identity. Check that local loops are not multiplied by a library or message redelivery policy that the dashboard cannot see.

Observe the total elapsed time and outbound attempt count for each case. The request should stop within the declared end-to-end budget, including queue waits where the task contract includes them. Test cancellation while an attempt is in flight. Verify that cancellation does not quietly restart the task through a different worker.

An external write needs its own retry contract. A provider timeout says nothing conclusive about an update sent earlier to another system. Keep the uncertain effect in reconciliation, retain its operation identity and inspect the receiving stub. Do not give a model retry permission to repeat an already submitted business command.

5. Verify containment without breaking status access

Owner: operations owner. Output: containment timeline. Trigger the sustained failure fixture and observe when new calls stop or become bounded. Check that the affected scope matches the dependency boundary. An unavailable generation route should not unnecessarily prevent the user from reading task status, cancelling never-started work or retrieving an existing accepted record.

The Circuit Breaker pattern separates blocked calls from limited recovery probes. That is a dependency-protection mechanism, not a business-completion decision. Your application's containment state must retain the information needed to choose what happens to each task.

Run fresh intake during containment. Verify admission limits, queue age and the response given when the queue cannot accept more work. A full queue should not silently discard a request after showing a successful receipt. A deliberate refusal with a clear status is better than an invisible promise the system cannot fulfill.

Inspect shared-resource use while faults persist. Status lookup and cancellation may compete with failing generation attempts for threads, connections or storage. Record which resource limited the service first. If containment protects the provider but exhausts the application's own workers, the drill has found an unresolved failure mode.

6. Exercise fallback only within its evaluated contract

Owner: domain acceptance owner and security reviewer. Output: fallback eligibility evidence. Select fixtures from the approved substitute's task scope. Verify the source permissions, data destination, response schema, evidence requirements and any supported tool boundaries. A plausible-looking substitute response is insufficient if it changes the meaning of an accepted task.

Use negative fixtures as well. A task outside the substitute's scope should remain held or unavailable, even when the alternate endpoint is healthy. Restricted inputs should not be forwarded to an unapproved data path. A tool the substitute cannot represent reliably should remain disabled rather than approximated through natural-language instructions.

Compare output under the same acceptance rubric used for the primary route. Record configuration revisions and any deliberately narrower capability. Test the user-facing explanation of that limitation. Avoid an interface that claims normal service while quietly removing required citations, approval checks or supported actions.

If the application has no approved fallback, test that fact explicitly. The exercise can pass by preserving tasks and reporting unavailability correctly. It does not need a second provider to look sophisticated. Use fallback output contracts when deciding whether a future substitute is worth evaluating.

7. Test the requester experience without imagined staffing

Owner: application owner with an independent observer. Output: captured status transitions. Read the interface as a requester, not as an engineer watching a trace. Does it say what is unavailable, whether intake was retained and what can happen next? A person should not need internal state names to understand whether their business request completed.

Capture messages for never-started work, interrupted generation and an uncertain external action. These must not share a misleading “failed, try again” instruction. In the uncertain case, test whether submitting again reveals the original pending task or creates another effect. Include a browser refresh and a repeated client request.

Do not offer live human assistance unless the organization actually staffs it. A bounded self-service status path, later response window or explicit inability to proceed can be honest alternatives. If a notification channel is optional, test that refusing it does not remove access to task status or the available recovery choice.

Verify cancellation wording through actual state readback. The application may stop future preparation while an earlier command remains unresolved. It should explain that distinction. Preserve evidence of what stopped, what already happened and what is still being checked instead of erasing the task to make the interface look settled.

8. Reconcile effects before choosing what may resume

Owner: recovery engineer with the domain owner. Output: per-effect recovery register. Restore access to the receiving stub while generation remains contained. Follow each simulated write by its stable operation identity. Classify its outcome as confirmed applied, confirmed rejected or unresolved. A missing receipt alone is not proof that nothing happened.

Use a case where the stub applies an update but drops the reply. Observe whether the application retrieves the applied state and suppresses a duplicate command. Add a case where the update and its notification have different outcomes. Reconciliation must cover each effect, not only the most visible record update.

Keep unresolved effects held if the receiving system cannot establish their outcome. Record the decision owner and next evidence required. Do not convert the record to a fresh operation simply to bypass a deduplication conflict. The exercise should prove that recovery can stop with uncertainty instead of manufacturing success.

For work that truly never started, recheck current inputs, permissions, cancellation and approvals before scheduling another attempt. If any changed, choose the appropriate new decision path. Use the action-recovery playbook for the detailed write-recovery contract; this outage drill tests that the contract survives dependency failure.

9. Calculate a drain plan with new arrivals included

Owner: operations owner. Output: bounded backlog recovery plan. Count valid queued tasks and their age, then measure sustainable accepted processing capacity under the recovery configuration. Include provider limits, review capacity and downstream handling. The fastest observed burst is not necessarily the rate the entire workflow can sustain.

Google's SRE chapter on handling overload explains why resource demand and retry behavior matter beyond a simple request count. Use that distinction when inspecting your own task mix. A short answer and a multi-stage investigation may consume very different capacity even when both appear as one queued item.

For a synthetic rehearsal, suppose a one-hour interruption leaves 60 valid tasks queued. Verified accepted capacity is 90 tasks per hour and new arrivals continue at 60 per hour. Net drain is 30 tasks per hour, so clearing that backlog takes roughly two hours if rates and task complexity remain stable. It is not cleared in forty minutes by dividing by gross capacity.

If net drain is zero or negative, state that the backlog will not clear under current conditions. Adjust admission, allocate tested capacity or change the service commitment through the owner. Do not promise a recovery time based only on endpoint health. Separately classify expired or cancelled entries so removing invalid work is not reported as successful completion.

10. Reopen gradually and test a second failure

Owner: exercise controller and operations owner. Output: recovery transition evidence. End the injected failure and permit only the declared probes. Check actual task outcomes, not merely successful transport responses. Restore fresh work and eligible backlog in controlled increments under the agreed observation windows and abort conditions.

Keep containment available during reopening. Reintroduce a failure while the backlog is draining. Verify that work already completed is not resubmitted, pending attempts retain their identity and admission returns to the correct degraded policy. A drill that only tests a clean return can miss the failure caused by oscillating dependency health.

Observe fairness and deadline behavior. New arrivals should not indefinitely starve older eligible tasks; older work should not monopolize capacity while urgent new tasks miss their contract. Define the policy with the domain owner. Record the actual selection order and the reason for any priority override.

If output quality, authorization checks or downstream reconciliation fails, stop expansion even if the provider error rate is low. The rollback is a return to the tested safe operating mode, with preserved records for in-flight tasks. It is not deletion of the queue or replacement of uncertain operations with newly generated identifiers.

11. Close every fixture with evidence and a remediation owner

Owner: independent observer with the application owner. Output: acceptance record. Compare observations to the expectations written before injection. Give each fixture a pass, fail or unresolved disposition with its evidence reference. Unresolved means an acceptance condition was not demonstrated; it should not be grouped with passing cases in the summary.

Use this review checklist. For each item, retain the fixture identity, observation time, responsible owner and evidence location:

  • Isolation excludes live business effects and the abort control was demonstrated.
  • Error classification distinguishes repairable application faults from dependency failure.
  • Total retries, elapsed time and recovery probes stay inside declared budgets.
  • Queue admission, expiry and cancellation produce truthful requester status.
  • Only evaluated task families and approved data paths use fallback.
  • Uncertain effects remain held until authoritative reconciliation permits a decision.
  • Replay rechecks current authority, input revision and task validity.
  • Backlog estimates use net capacity and disclose their assumptions.
  • A second failure during reopening returns to a safe state without duplicate effects.

Assign each failed or unresolved condition a concrete repair, owner and retest fixture. Separate a documentation correction from a code change and from a missing product decision. Keep the original failed evidence. An updated runbook does not demonstrate that the implementation now behaves differently.

The next action is an owned retest or a bounded release recommendation, not an automatic production drill. Use the AI change-release playbook before exposing new recovery behavior. For help applying the findings, explore production AI systems or share your outage question. Reading and downloading this guidance does not require contact details; an enquiry is a separate optional choice.