Test a Shared Dependency Without Losing the Recovery Path
Map customer journeys onto shared dependencies, rehearse one bounded failure and reconcile tenant impact, retries and unfinished work before accepting recovery.
trigger="Several customer journeys use one platform dependency, but the team has not demonstrated what happens when it slows down or becomes unavailable." owner="The platform owner accountable for the approved target boundary and recovery record." participants={['Service maintainer', 'Exercise operator', 'Independent observer', 'Security reviewer', 'Business-operation owner']} prerequisites={['An isolated environment with synthetic tenants', 'An explicit dependency and target inventory', 'Independent stop and recovery access', 'Agreed impact limits and evidence handling']} outputs={['A journey-to-dependency map', 'An approved fault contract', 'Per-tenant baseline and fault observations', 'A pending-work reconciliation', 'An acceptance record with limitations']} doneWhen={['The injected condition is removed', 'Every in-scope journey has fresh recovery evidence', 'Pending operations have an owned disposition', 'No unapproved targets or external effects occurred', 'Unexercised paths remain explicit']} />
Rehearse the customer operations that share a dependency
Rehearse a shared dependency failure by identifying which customer operations need it, defining their permitted degraded behavior and testing one bounded fault in an isolated environment. Measure each tenant journey separately. Keep the stop mechanism and evidence collection outside the targeted failure path, then reconcile outstanding work before declaring recovery.
A dependency can be healthy while a customer remains unable to begin a new session, a worker replays an old operation, or one tenant waits behind another tenant's backlog. Check the affected operation against its contract during the fault and after restoration. The resulting evidence covers those operations and conditions; it cannot certify availability for the whole platform.
The example throughout is hypothetical: two synthetic tenants share a permission service in a test environment. Existing bounded reads and new privileged writes use different authorization paths. The exercise must not treat a failed permission check as permission to continue. An independent observer records outcomes, and a recovery identity does not rely on the targeted permission endpoint. Replace these assumptions with your actual dependencies before executing anything. This procedure is not authorization to interrupt a live system.
1. Name the journey and the decision it must preserve
Owner: business-operation owner with service maintainer. Output: journey contracts. Pick a small set of operations with different dependency behavior. For example, reading an already-authorized record, opening a new session and submitting a privileged update. Specify which result is allowed if the dependency cannot answer: refuse the operation, serve explicitly permitted bounded data, or retain a request for later review. Define the customer-visible wording and how the caller can determine whether an operation completed.
Separate safety from convenience. An authorization outage must not widen access. A write that may already have committed must not be presented as safely retryable merely because its response timed out. Define the authoritative record used to settle that uncertainty. If a read may use previously established permissions, record the validity boundary, revocation requirements and allowed data scope rather than vaguely calling the behavior a cache fallback.
Choose acceptance thresholds from the operation's actual agreement and risk tolerance. Record the measurement window, sample count and excluded traffic. An invented target such as a universal five-second recovery time would conceal rather than resolve the decision. If the operation has no accepted degraded contract yet, that is design work to finish before fault injection.
2. Trace ordinary use, startup and recovery dependencies
Owner: platform owner with application maintainers. Output: dependency register. Trace the chosen journeys through shared identity, name resolution, configuration, secrets, queues, storage and telemetry where they are actually used. Include nested calls, client libraries and workers that run after the initial response. Record tenant ownership, resource identifiers, endpoints, credential boundaries and the evidence supporting each relationship. An architecture diagram is a starting hypothesis until configuration and observations support it.
Distinguish a dependency used for every request from one needed to establish new access, replace a process or apply configuration. AWS's control-plane and data-plane explanation describes administrative operations separately from the primary service function. Use that distinction to ask which transitions your workload requires. It does not establish that your application's existing requests will survive a provider or platform outage.
Trace the recovery operator's access too. A console, tunnel, token refresh or monitoring login may need the same failed service. Mark those paths explicitly and prepare approved independent access where appropriate. Do not weaken normal authentication as a substitute. If recovery depends on an untested emergency identity, the exercise cannot claim to have demonstrated an operable stop path.
3. Separate the fault boundary from the evidence boundary
Owner: exercise operator with independent observer. Output: scoped exercise map. Select the narrowest fault capable of answering the journey question. For a first rehearsal, prefer an isolated dependency substitute or a dedicated test endpoint that can return controlled delay or errors. A substitute establishes client behavior for that condition, not the real service's internal failure behavior. Keep that limitation in the report.
The diagram shows a proposed test arrangement, not Ampity's production topology. Solid arrows mean a request path. The fault is applied to the isolated shared permission endpoint only. The observer and stop access are intentionally separate; their independence must be checked in the actual environment. The existing read path remains subject to its own permission validity contract and does not bypass authorization.
Place external delivery systems behind inert test receivers. A delayed synthetic write must not become a real email, payment, user update or customer notification when connectivity returns. Verify outbound credentials and destinations, not just the environment label. Unknown shared resources are a reason to hold the exercise, not a reason to assume that staging is isolated.
4. Write a fault contract before choosing the tool
Owner: exercise operator with security reviewer. Output: approved fault contract. Record the condition, target identifiers, start method, maximum duration, maximum synthetic load, affected tenants and expected downstream effects. Distinguish connection failure, a slow response, a rejected request and a successful response with invalid content. These are different tests. Begin with one condition so the resulting behavior can be attributed to a known cause.
Specify exactly how the fault is removed and what persists afterward. A latency rule may expire, but a killed process needs replacement and a changed configuration needs restoration. Save the accepted configuration revision and verify that the reversal procedure restores that revision or a separately approved successor. Do not promise that every chaos tool undoes every action automatically.
AWS FIS's stop-experiment documentation says outstanding post actions complete before an experiment stops and that a stopped experiment cannot resume. Inspect the chosen action's own behavior before using it. A tool reporting stopped is evidence about the experiment state, not proof that customer work, configuration or retained data has returned to the accepted state. No FIS action is required for the isolated substitute used in this example.
5. Establish a per-tenant baseline and workload denominator
Owner: independent observer with service maintainer. Output: baseline evidence. Run the synthetic journeys without a fault using known fixture identities. Record attempts, accepted operations, refused operations, completed effects, uncertain effects, latency distribution and workload volume for each tenant. Include the request identity and relevant service revision. Avoid collecting confidential payloads when a test identifier and disposition are sufficient.
Track an operation separately from its attempts. Three retries of one submission are not three business requests, and an accepted response is not necessarily a completed background effect. Use an operation ledger to join the initial request, queue receipt and authoritative completion record. Record how you detect duplicate effects and how long completion may legitimately take. Otherwise a shorter observation window can produce an artificial improvement or an unexplained backlog.
Capture the low-traffic tenant as carefully as the high-traffic tenant. An aggregate error rate can hide a complete outage for the smaller population. Keep the denominator and observation coverage visible in each result. If a journey generated no attempts, report it as unobserved rather than healthy. Baseline instability or missing identity joins prevents useful comparison and should postpone the fault run.
6. Prove stop conditions and observation loss handling
Owner: independent observer with exercise operator. Output: tested stop evidence. Define abort criteria for scope escape, prohibited authorization behavior, unexpected external effects, excessive pressure and loss of evidence. Test the stop mechanism with a harmless rehearsal before introducing the real fixture condition. Confirm who can invoke it, from which access path, and how the operator verifies that the fault is no longer being applied.
For an AWS FIS experiment, stop conditions use CloudWatch alarms to stop an experiment at a defined threshold. Do not translate that into an assumption of instantaneous containment. Account for observation, evaluation and removal delays in the selected action and monitoring arrangement. The exercise must remain bounded while those mechanisms act, with a separately available manual stop path.
Test missing observations deliberately. CloudWatch's missing-data documentation explains that configured treatments can differ, including retaining the current state or treating missing data as good or bad. Choose behavior for the actual metric and alarm type. In this proposed procedure, loss of the required journey evidence is an independent reason to stop and investigate, even when a platform dashboard still looks acceptable. Never infer success solely from the absence of an error event.
7. Apply one fault and account for every selected target
Owner: exercise operator. Output: target readback and fault timeline. Immediately before the run, resolve the intended targets and compare them with the approved register. Record the selected environment, endpoint and tenant fixtures. Make the test unable to use production credentials or discover additional resources through broad selectors. A correct template reviewed yesterday is not a readback of what will be affected now.
Start the baseline workload, record the actual fault-start evidence and introduce the approved condition. Observe one small bounded interval before any increase. Keep background jobs, scheduled consumers and load-generator retries within the same total request budget. If the injector fails to apply the condition, label the run invalid rather than crediting the platform for surviving an outage that never occurred.
Do not modify application configuration mid-run to make the result look better. If a correction is necessary, stop, preserve the result and create a new run with the changed revision. This separates diagnosis from validation. Record early stops as valid safety outcomes with limited behavioral evidence, not failed paperwork to be omitted from the final report.
8. Inspect retry pressure and tenant interference
Owner: service maintainer with platform owner. Output: pressure and isolation findings. Measure downstream attempts per original operation, concurrent outstanding calls, pool waiting, retry timing and queue growth. Identify retry behavior in the browser or caller, gateway, service and worker where those layers exist. Adding the same local retry policy at every layer can multiply pressure; do not evaluate one library in isolation from the request path.
Observe what happens to the second tenant when the first tenant encounters repeated failure. Shared connection pools, worker limits and recovery queues may transmit pressure even when data access is correctly separated. Compare per-tenant outcomes with the baseline, while retaining aggregate resource measurements. A tenant-specific prefix or access rule is not evidence of isolated execution capacity.
For AI-assisted operations, include model retries and tool attempts in this register. A model switching providers does not repair a failed permission decision or give a tool broader authority. Keep privileged writes unavailable when their required decision cannot be established. Any useful read-only degraded answer needs its own data freshness, permission and disclosure contract. This drill tests that boundary rather than making AI a general recovery mechanism.
9. Separate refusal, pending work and completed effects
Owner: business-operation owner with service maintainer. Output: operation disposition ledger. Classify every synthetic operation using authoritative evidence. A refusal before submission, an accepted queue item, a committed write whose response was lost and an unresolved request require different handling. Record the tenant, operation identity, attempted revisions, known receipts, actual effects and next allowed action. Avoid a single success/failure label that encourages blind resubmission.
Reconcile both directions: every accepted operation must have a disposition, and every observed effect must map to an authorized operation. A hidden duplicate can coexist with apparently successful requests. Inspect requests initiated before the fault as well as those accepted during it, because they may cross the boundary after their original permission or caller deadline changes.
Do not automatically delete a pending item to make the queue look clean. Determine whether it should finish, be cancelled, require renewed authorization or remain on hold. Record who owns the decision. Where the test cannot inspect a downstream effect, retain the uncertainty and its owner. The same discipline applies to AI tool writes, automated notifications and ordinary application transactions.
10. Remove the fault and verify the recovery transition
Owner: exercise operator with independent observer. Output: reversal and fresh-journey evidence. Remove the injected condition and read back the target's configuration and current state. Record when removal was requested and when observations establish that it took effect. Then run fresh operations with new identities. Reusing only existing connections or previously authorized sessions may conceal a broken startup or credential-refresh dependency.
Test controlled backlog resumption separately from normal new traffic. Limit replay according to the approved capacity budget, preserve operation identities and inspect renewed authorization requirements. Watch whether recovery pressure delays the second tenant or repeatedly reconnects to the restored service. Restoration of the dependency is only the beginning of this observation window.
An early abort follows the same recovery path. Stop admission of new synthetic work, remove the fault, preserve the ledger and settle already-started work. Use the independent recovery access if the ordinary administration path is unavailable. Do not resume the experiment or replay uncertain operations until a new decision approves the scope. The report must distinguish fault removal time, journey recovery time and pending-work disposition time.
11. Review the acceptance checklist against actual evidence
Owner: independent reviewer. Output: acceptance record. Use these acceptance criteria as a review checklist, not a score calculated from average uptime. Attach evidence references and record accepted, rejected or unknown for each item. A missing record remains unknown. The hypothetical permission-service run passes only the tested conditions, workload and revisions; it does not establish behavior for a database outage or a different tenant population.
- Scope: Selected targets match the approved register, and no external effect or unapproved resource was reached.
- Safety: The stop path was exercised, observation loss had a defined response, and permission failure never widened access.
- Tenant impact: Each selected journey has its own baseline, fault and recovery denominator, including refusals and unobserved cases.
- Attempt pressure: Retry and concurrency evidence covers every participating layer and the second tenant's shared-resource exposure.
- Operation settlement: Accepted work and observed effects reconcile by identity; uncertain work has an owner and permitted next action.
- Recovery: The fault is removed, fresh access and startup paths work, and controlled backlog processing remains within the accepted limits.
Record disagreements between dashboard data, client observations and the operation ledger. Resolve the mismatch or explicitly limit the conclusion. An acceptance signature should reference a specific run and evidence package, not silently turn a partial rehearsal into approval for a larger production experiment.
12. Turn findings into a bounded engineering change
Owner: platform owner with service maintainer. Output: remediation and retest plan. Choose the smallest change supported by the observed failure. Examples include a client deadline that actually covers the whole call, a retry budget shared across layers, a separate worker allocation or a safer customer-visible pending state. These are possible responses, not universal recommendations. Keep the original rejected run so the next test can establish whether the change addressed its specific defect.
Retest using the same identities and workload shape where comparison requires them, while generating new operation identifiers for every new effect. Record changed dependencies, policies and software revisions. Repeat the stop test if the change affects access or monitoring. If the environment cannot reproduce a relevant production relationship, report that gap and define a separately approved validation step rather than manufacturing a passing result.
Start the next review with one completed dependency register and disposition ledger. The control-plane outage article helps distinguish running traffic from recovery transitions. The shared-cache isolation article examines another common source of tenant interference. Bring the observed boundary to a reliability review or DevOps and SRE review when the next change needs accountable implementation. Reading and downloading this procedure does not require sharing your email.