Shared Platform Dependencies: Choosing a Defensible Blast Radius
Compare shared services, bounded resource pools and independent cells against explicit failure scenarios, tenant impact, authorization and recovery evidence.
audience="Teams whose customer workflows depend on shared identity, configuration, queues, data stores or AI services." decision="Choose which failure must be contained, which customer functions may degrade and which dependencies need stronger capacity or state boundaries." position="Keep sharing where it is justified, but test the actual propagation path. Separate resource limits from state isolation and preserve current authorization and uncertain business effects during recovery." scope="A proposed engineering comparison with synthetic tenants and fixtures, not a measured Ampity deployment, cloud availability guarantee or instruction to run a live fault experiment." outputs={['A customer-function contract', 'A dependency and fault register', 'An isolation-options decision', 'A propagation and admission review', 'Independent tenant-impact fixtures', 'An owned migration and recovery record']} />
Executive summary
A shared platform service can make a product cheaper to run and easier to govern. It can also turn one slow tenant, invalid configuration or unavailable permission check into a problem for customers that did not cause it. The useful question is not whether shared services are good or bad. It is which failures can travel through the shared boundary, how much customer work they can affect and what evidence would justify a more expensive separation.
Compare three choices deliberately. A shared service with admission controls may be sufficient when the dominant risk is one caller exhausting bounded capacity. Separate resource pools can protect work classes while preserving common code and state. Independent cells can contain selected state, deployment and capacity failures, but they require partitioning, routing, migration and operating discipline. Separate containers that still depend on one database do not establish independent cells for a database failure.
Make the failure claim specific. A proposed boundary may contain a tenant's expensive query without containing a bad release applied everywhere, a regional outage or a central credential failure. Describe the fault, affected customer function, duration observed and evidence gaps. A healthy global average does not prove that a smaller tenant survived. A successful health check does not prove that delayed customer work completed or that its security conditions remained valid.
This paper uses a hypothetical document-workflow platform with tenants A and B. It includes document viewing, approval submission and an optional AI explanation. The diagrams and fixture expectations are proposed designs, not customer results or production experiments. Official sources support particular design principles. The workload contract, thresholds and acceptance decisions belong to the responsible organization. No pattern here guarantees an availability level or permits weakening access control to make an incident dashboard green.
Define the customer function before the isolation boundary
Start with a function a customer recognizes. Viewing an already available document, submitting a new approval and requesting an AI explanation do not have identical dependencies or consequences. A platform may continue serving existing documents while blocking new approvals because it cannot establish current authority. Describing both as the website is up hides the difference. Write a separate success, refusal, delay and recovery contract for each function before selecting a deployment topology.
For the proposed example, an authorized document view must apply the declared access policy. An approval submission must not be reported complete without evidence of the recorded effect. The optional AI explanation can be unavailable while ordinary document work continues. These are fixture choices, not universal rules. Another product may make AI interpretation mandatory or prohibit any stale access information. Its acceptable degraded state would therefore differ even if it uses the same infrastructure components.
Record what the user sees when the contract cannot be met. A clear pending state is different from a fabricated success, and a deliberate refusal is different from an internal error. Where a request is accepted for later processing, describe the supported status lookup and cancellation semantics. Where the operation cannot be admitted, avoid presenting it as queued. The customer contract needs to distinguish work that never began from work whose final effect is currently unknown.
Identify the owner who can accept a reduced function, not just the engineer who can implement it. Product, security, data and operations owners may need to agree on different consequences. Keep the approved boundary attached to a concrete workflow and configuration revision. A general incident policy saying use cached data when possible is too broad to decide whether an old document permission remains acceptable after access withdrawal. The review must name the permitted freshness and withdrawal behaviour.
Declare the fault model and the limits of the claim
List the failures being considered separately: slow responses, explicit unavailability, malformed replies, exhausted connections, corrupt shared configuration, rejected credentials, bad releases and lost destination acknowledgements. Different controls address different cases. A timeout bounds waiting but does not establish that an attempted write failed. A second instance can survive a process crash while accepting the same corrupt configuration. A separate queue can protect admission capacity while leaving the consumers dependent on one unavailable store.
For each fault, record its injection point, scope, duration, load conditions and expected observation. A permission-service timeout in tenant A's test path is not equivalent to removing the entire shared permission service. Likewise, one tenant's heavy background task is not equivalent to regional network loss. Keep the conclusion at the level exercised. If the test environment does not represent a common database limit or account quota, disclose that gap rather than promoting the result into a broad isolation claim.
Distinguish a capacity fault from a semantic fault. A service can respond quickly with an incorrect authorization result or incompatible schema. Capacity limits might keep other callers responsive while the wrong decision still reaches them. Conversely, rejecting an invalid response can preserve correctness while reducing availability. The review should not consider correct refusal a security failure merely because its response is unsuccessful. Report contract correctness and timely useful work as separate observations.
Include the observation and recovery paths in the fault model. If the experiment also removes the operator's login, stop control or evidence store, the team may lose the ability to understand or reverse it. A claimed containment boundary needs independent access appropriate to the scenario. This paper does not authorize live injection. First rehearse with inert dependencies and synthetic records in a bounded environment, then obtain separate operational approval for any consequential test.
Trace dependencies through serving, refresh and recovery
Construct a dependency register for each customer function. Record the caller, callee, synchronous or asynchronous relationship, timeout, retry owner, resource pool, state location and authority requirement. Follow indirect paths: a document service might call a policy service that queries a central database and refreshes credentials from another system. A picture showing only document service to policy service misses the deeper common failure. Keep the detailed register beside the visual instead of compressing every relationship into tiny labels.
Trace the path over time as well as per request. An existing process may serve a validated configuration until its refresh deadline, then become dependent on the configuration service. A worker may continue until its credential expires. A recovery process may need to launch new capacity through a different control path. Record the time-dependent transition and what the contract permits then. An observation made before refresh cannot establish that the same function remains available after it.
Review where resource consumption crosses the boundary. Waiting for a slow dependency may occupy a frontend connection, an application worker and a database connection at once. An optional task can therefore consume capacity needed by the essential path even when its own code is separate. Track concurrency, memory, connection pools and worker scheduling, not only network calls. A dependency absent from the request's output can still determine whether the request obtains a worker.
Keep ownership and evidence gaps explicit. The platform team may own the client adapter while another team or provider owns the service. Document which behaviour is under local control and which is a contract to verify. An undocumented provider limit is not infinite capacity; a missing trace is not proof that no dependency exists. Use representative runtime evidence, configuration readback and maintenance knowledge together, while recording the environment and paths that were not inspected.
Compare shared capacity, bounded pools and independent cells
Keeping one shared service is a valid option when its operating envelope, correctness contract and recovery path meet the workload's needs. It avoids duplicated data, multiple upgrade paths and routing complexity. Add explicit admission and caller controls rather than assuming aggregate autoscaling will protect every tenant. Its limitation remains the shared service and state: a defect in their common logic or a full outage can affect all callers that require that capability.
Bounded pools separate a resource such as worker concurrency, connections or queue capacity. They can distinguish interactive viewing from expensive AI tasks, or protect one tenant's reserved capacity from another's burst. This is usually a narrower change than duplicating the whole application. It is not full state isolation. If both pools depend on the same failing database, central policy logic or credential, that shared failure remains. Name the protected resource and the remaining common dependencies in the decision.
AWS's bulkhead guidance describes cells with independent state, a bounded size, persistent partition mapping, minimized cross-cell interaction and staggered releases. It also identifies the router as a shared layer needing a simple design. Use these as properties to evaluate, not a label to apply to any multi-instance deployment. A cell proposal should say what state, capacity and releases are actually independent for the declared fault.
Independent cells introduce costs that are easy to omit from a topology slide. The team must place tenants, monitor skew, move growing workloads, reconcile migrations and support multiple operating instances. Shared analytics or global reporting can reintroduce runtime dependencies if implemented synchronously. Compare these costs against the consequence of the specific failure, rather than assuming more isolation always improves the business result. Choose the least complex option that satisfies the accepted contract, while preserving a path to strengthen the boundary when evidence changes.
Read the topology against the fault being contained
The first visual compares a common dependency with a proposed cell-local replacement. It answers where state and serving capacity are shared, not how every application feature works. Requests have a stable placement into one cell; each cell owns the shown service and state. The shared routing layer remains outside that containment boundary. The diagram deliberately uses neutral component names because no particular vendor deployment has been selected or verified.
Do not interpret the second option as a promise that tenant B is unaffected by every tenant A incident. The claim is limited to a selected cell-local service or state fault, under verified placement and capacity conditions. A shared router defect, simultaneous release, common credential or regional impairment can still cross the proposed boundary. Those faults need separate reviews. Adding a dashed cell frame does not remove them from the actual execution path.
The migration decision should preserve the distinctions omitted from this overview. Identify the real partition key, data placement, policy distribution, tenant lifecycle and operator access. If a document shared by two tenants requires cross-cell access, either redesign that operation or record the dependency it creates. A genuine shared business object may constrain isolation more than the preferred layout suggests. The architecture should follow the business ownership model, not invent independence by hiding those relationships.
Bound admission using the resource that actually becomes scarce
Google SRE's overload discussion explains why requests per second alone can be a poor capacity proxy when requests consume different resources. The local implication is to measure the constrained resource and work mix before setting admission limits. A document preview, large export and AI explanation may use very different memory, connection and processing budgets. Equal request counts do not imply equal load or fair access to useful service.
For the proposed platform, define separate concurrency controls for document viewing and optional explanation. Identify where a waiting AI request holds frontend or worker capacity. If a shared worker pool has no reserved path for essential viewing, limiting calls to the AI provider may leave the frontend exhausted before the limit takes effect. Place the control before the costly allocation and verify what happens to rejected work. A limit should protect a named resource, not merely create an additional metric.
Choose fairness and priority deliberately. Per-tenant limits can reduce noisy-neighbour effects, but a legitimate high-volume customer may require more capacity than a small one. A product can allocate an agreed entitlement, isolate expensive work classes or use a bounded scheduling policy. Record starvation risks and exceptions. Do not quietly treat all errors for a lower-priority tenant as acceptable simply because a global success measure improved. Every supported tenant still needs a declared service contract.
Test capacity with representative payloads and impaired conditions, not only a healthy average. Verify that the reserved essential path remains useful when optional tasks are slow, canceled or rejected. Record admitted, rejected, expired and completed work separately, plus the resource evidence explaining each. Derive thresholds from measured behaviour and owner-approved objectives. This paper does not provide universal concurrency numbers, because copying a limit from another workload can either waste capacity or fail to contain its actual bottleneck.
Keep retries from exporting a local fault
AWS's retry guidance recommends bounded retries, backoff and jitter, with attention to idempotency and retry multiplication across layers. Inspect the actual SDK and adapter behaviour rather than adding another loop on top. The boundary review should establish which layer owns retries, which errors are eligible and which end-to-end deadline stops further attempts. A caller timeout must not leave uncontrolled background retries consuming shared capacity indefinitely.
Use a small synthetic arithmetic check before measuring the system. If three nested layers each make up to two total attempts, one logical request can generate up to eight attempts at the deepest dependency when every path retries. This is an illustrative upper bound under the stated assumptions, not a production rate or guarantee. The point is to count actual attempts across the stack. Changing only the outer retry setting can leave amplification in a proxy or SDK unchanged.
Separate harmless read retries from effects that might already have committed. A permission lookup can be retried under its freshness contract; an approval write requires an operation identity and supported reconciliation semantics. Reusing an idempotency mechanism is meaningful only when the destination's scope and retention match the operation. Do not infer that a timeout means absence of effect or that a new worker may safely create a second approval. Preserve uncertain effect state until trusted evidence supports a disposition.
Review stop conditions at each boundary. Cancellation of an incoming request may not cancel an outbound provider task, and opening a local breaker may stop new calls without removing already admitted work. Measure residual operations and capacity after the stop, rather than reporting containment from the control changing state. Keep the request's deadline, retry budget and business deadline distinct. A technically successful late result can still be unusable for the customer's time-sensitive workflow.
Use queue age and disposition, not just queue depth
A queue can decouple immediate admission from processing, but it also stores obligations. Define what the producer promises when it accepts a task: processing by a declared deadline, best-effort completion or a status that may later expire. A durable queue does not make every task useful forever. Track age, expiry, retries and tenant ownership alongside depth. A short queue containing one very expensive or stale task can be more consequential than a longer queue of cheap current work.
The second visual shows a failure propagation path in reverse of a common dashboard narrative. A dependency slowdown retains active work; retries add demand; shared worker capacity can then be consumed by waiting; essential customer functions lose capacity. The controls are attached to their actual boundaries. Admission limits, finite waiting and a separate essential pool address different points. The diagram is an illustrative resource relationship, not evidence of a measured incident or a universal sequence.
For recovery, classify queued work before draining it. A task can be current and permitted, past its business deadline, canceled, based on superseded inputs or associated with an uncertain destination effect. Restoring the dependency does not make all those tasks eligible. Resolve each according to its contract, preserving a supported terminal or review state. Avoid sending an entire accumulated backlog into a recovering service without checking whether its capacity and retry limits can tolerate the surge.
Degrade optional functions without relaxing authority
AWS's graceful-degradation guidance ties reduced function to the core business requirement and requires failure paths to be tested. Its examples are not permission to cache every security decision or substitute arbitrary data. For the proposed platform, optional AI explanation can be unavailable while permitted document viewing continues. The interface should identify the unavailable feature without inventing an explanation or promising it will arrive when no task was accepted.
Authorization is different. If the service cannot establish access under the declared policy, the proposed fixture refuses the protected operation. It does not bypass the check to preserve availability. A locally available policy snapshot is usable only under an explicit validity and withdrawal contract. Test access changes during impairment, including a previously authorized user whose permission was withdrawn. Stale access should not be described as graceful degradation merely because the document still opens quickly.
Keep the fallback simpler than the main workflow. A second model provider, substitute database or alternate region introduces its own access, data, schema and effect semantics. It may share the same constrained network or credentials. Disabling an optional stage can be safer and easier to explain than invoking an unreviewed alternate path. If an alternate path is necessary, evaluate it as a separate implementation with independent failure and recovery fixtures, not as an automatic improvement in resilience.
Make reduced function visible in evidence and communication. Record whether the user received the core result, a partial result, a justified refusal or an accepted pending task. Avoid marking a response successful only because its HTTP status is two hundred. Product and security owners should approve these distinctions before an incident. A clear refusal can preserve the service's correctness contract even while availability is reduced; the team should report both facts rather than collapsing them into one green status.
Separate serving stability from changes needed to stay alive
The Amazon Builders' Library discussion of static stability distinguishes ongoing serving from control-plane changes and illustrates preparing capacity before an impairment. Apply the question to your own application dependencies as well: can existing useful work continue without fetching a new configuration, launching replacements or obtaining refreshed credentials? This does not mean every workload should overprovision identically or that existing authority may remain valid without its required checks.
Record the survival window of the actual serving path. An existing document worker may have valid configuration and capacity now, but a refresh requirement, credential expiry or restart may change its dependencies. Test both before and after that transition. New tenant creation and tenant placement changes can be held while existing work continues, provided the product contract permits that state. A control-plane label alone does not prove the serving path is independent for the required duration.
Keep validated local state distinct from missing or corrupt state. A failed refresh must not automatically replace a known configuration with an empty default. Equally, retaining an old snapshot is not correct when mandatory withdrawal requires it to stop. Record snapshot version, allowed use and the owner's invalidation policy. A fallback that preserves yesterday's functionality while ignoring today's access restriction violates the intended contract even if it looks statically stable in a performance test.
Review replacement capacity and operator actions as separate paths. If every recovery attempt needs a central service that is currently unavailable, the system may survive its initial impairment but fail when a worker is lost. Verify that observation and stop controls remain accessible in the approved scenario. Describe which transitions are deliberately unavailable, which are prepared in advance and which remain untested. Do not turn a successful short serving test into an indefinite continuity guarantee.
Inspect common causes that a cell diagram cannot remove
Several nominally separate cells can share a router, release pipeline, secrets source, account quota, network dependency or operating team. Some sharing is unavoidable or justified. The review must disclose its effect on the claimed failure boundary. A bad release promoted to all cells at once can produce simultaneous failures without any runtime call between them. Independent storage does not contain that deployment path. Staggered rollout and evidence from the first affected cell address a different risk than storage partitioning.
Google SRE's cascading-failure discussion describes how losing capacity and shifting traffic can overload remaining healthy service. The proposed design therefore must not reroute all of tenant A's impaired-cell demand into B merely because B is healthy. Verify spare capacity and eligible data placement before any transfer. A failure boundary can be defeated by recovery automation that exports the load it was intended to contain.
Assess central policy and telemetry paths carefully. A shared policy distribution process can be outside ordinary serving but required at withdrawal time. A shared evidence store can support analysis without being a synchronous requirement for every read, unless the audit contract makes it mandatory. Record that actual distinction. Optional observability should not silently become a serving dependency through a blocking exporter, while required audit evidence should not be silently dropped to keep serving responsive.
Keep failure-domain claims layered. Tenant partitioning, process isolation, zone separation and regional recovery protect against different classes of impairment. None is a substitute for current authorization or semantic correctness. Identify which shared risks the organization knowingly accepts, the owner of each exception and the trigger for revisiting it. An architecture with disclosed limitations is more useful than a diagram that appears completely isolated because it omits everything inconvenient to the claim.
Measure tenant outcomes with independent expectations
Choose observations before running a rehearsal. For tenants A and B, record attempted customer functions, admitted work, correct completion, justified refusal, expiry, unknown effect and remaining pending work. Use per-tenant and per-function denominators. Tenant B's low traffic should not disappear inside A's large burst, and rejected requests should not be removed from the evaluation merely to improve apparent success. Missing observation is an evidence gap, not an unaffected tenant.
Measure the mechanisms needed to explain the outcome. Capture bounded concurrency, connection use, attempt counts, queue age and selected configuration versions with permitted operational evidence. A latency graph without the admission record cannot distinguish successful containment from users being silently dropped before measurement. Preserve enough correlation to reconcile the fixture, without recording protected document payloads or credentials in an unrestricted trace. Observation permission and diagnostic usefulness are separate design decisions.
Under the synthetic contract, essential viewing for tenant B must retain its accepted service behaviour when only tenant A's optional explanation pool is impaired. Tenant A's explanation should become explicitly unavailable or reach its declared pending disposition. If the shared authorization service itself fails and no currently valid decision is available, both tenants' protected operations are refused. That second result contradicts a broad all-failures containment claim, but it can satisfy the deliberately narrower security contract.
Vary the fixture without changing several causes at once. Test slow responses, invalid replies and common-state failure separately. Include a deliberately weak implementation that uses one unbounded pool for every work class; verify that the harness notices the lost essential capacity. A later retry completing the request does not erase the missed original deadline. Keep initial service outcome and eventual reconciliation in separate fields so recovery cannot retroactively make an impaired customer journey successful.
Use counterexamples to reject an overbroad isolation claim
The acceptance worksheet describes proposed outcomes under the preceding synthetic contract. It is not a test-results table. All effects stay inert and all recipients are fabricated. A passing result supports only the particular topology, work mix and fault exercised. The reviewer must still assess omitted shared dependencies and time-dependent transitions before approving wider use.
| Controlled fault | Expected customer disposition | Required evidence | | --- | --- | --- | | Tenant A optional explanation pool stalls | A explanation unavailable; B essential view remains permitted and useful | Per-function outcomes and independently bounded pools | | Common authorization service fails; no valid decision exists | Refuse protected operations for both tenants | Current authority gap and explicit refusal, not bypass | | Cell A local state fails in the proposed independent design | Hold A state-dependent work; B follows its unchanged contract | Actual state separation, placement and B denominator | | All cells receive the same invalid release | Reject a cell-local containment claim | Shared release revision and correlated impact | | Approval times out after a destination attempt | Keep its effect uncertain until supported reconciliation | Operation identity, destination readback and pending disposition |
Do not normalize these different cases into one requirement that B always succeeds. Some faults intentionally require refusal to preserve correctness. The test should reject a dependency bypass, unexplained mixed placement or unsupported completion even if the response is fast. Record which claimed boundary each counterexample challenges and which control is expected to detect it. That gives the reviewer a reasoned acceptance decision rather than a collection of screenshots labelled healthy.
Price and stage the migration without changing the business contract
An isolation proposal needs a cost and operating comparison, not just an infrastructure bill. Account for duplicated capacity and state, minimum idle footprint, backup coverage, telemetry, release management, tenancy migration and support effort. Compare useful accepted work under the chosen fault model, not cheaper requests achieved by rejecting the difficult ones. An optional AI pool can often be separated without duplicating the whole document platform. A common state defect may require a stronger and more consequential change.
Select a bounded migration unit and record authoritative placement. A tenant should not write into an old cell while reading an incompatible snapshot from a new one because routing changed independently of state. Define admission hold, data transfer, consistency checks, placement cutover and rollback eligibility. The design needs an owned answer for in-flight operations and delayed messages. A copied database is not proof that the customer's workflow may safely resume on it.
Rehearse with synthetic tenants first. Verify current authority, document ownership, status lookup, expected refusals and pending-work reconciliation after the move. Cross-cell business objects need explicit treatment. Keep a migration journal with observed receipts and unresolved effects rather than assuming all work stopped at the cutover timestamp. Reversal may require reconciliation, not just changing a routing entry, especially once new writes exist in the destination cell.
Choose acceptance and escalation triggers before the pilot. Evidence of one tenant exhausting common capacity may justify pool separation; recurring cell-local state or release risks may justify fuller isolation. Record which risk remains accepted and what would change that decision. Avoid building an elaborate cell platform before the team can operate its admission, placement and recovery contracts. The purpose is a defensible customer boundary, not the maximum number of boxes on an architecture page.
Reconcile recovery and close the decision with owned evidence
Restoring a dependency response is the beginning of recovery, not its completion. Identify work that was refused, accepted but unprocessed, expired, canceled or attempted with an unknown destination effect. Use supported destination evidence to reconcile the last group. Do not recreate an approval merely because the local record has no success acknowledgement. Resetting a queue, breaker or process cannot settle an external effect, and deleting uncertain records destroys the evidence needed to decide safely.
Drain eligible backlog within the recovering service's verified capacity and current authority contract. Reassess deadlines, source revisions and cancellation before execution. Monitor per-tenant results so a large backlog does not starve another customer's new work. Keep original deadline failures separate from eventual completion. The incident owner should state what is restored now, what is still pending and which effects remain unresolved instead of announcing recovery from one successful synthetic probe.
Use this review checklist before accepting a boundary: the customer functions and degraded states are owned; the dependency register includes serving, refresh and recovery; the fault claim is explicit; resource limits protect the actual scarce resource; state and common causes are disclosed; authority is not weakened; tenant denominators include refusals and unknowns; and migration and recovery have independently observed evidence. Reject the decision if any essential observation remains assumed rather than demonstrated.
Start the next review with one completed dependency register and one proposed fault. Use the shared-dependency rehearsal playbook for the executable procedure and control-plane outage discussion for time-dependent serving questions. Ampity's reliability review and DevOps and SRE services can help define an acceptance scope. The paper and its PDF can be used without submitting an enquiry; contact is optional.