What Keeps Running When the Control Plane Fails?
Review control-plane outage readiness by tracing serving, credential refresh, replacement capacity and recovery dependencies through customer functions.
Review the customer function before the plane label
Existing workloads may continue serving requests during a control-plane failure when their required runtime dependencies remain available. That does not establish that new capacity, configuration changes, credential renewal or recovery operations will succeed. Review each customer function and the period for which its dependencies remain usable, rather than giving the entire platform one availability label.
Consider a hypothetical report service. A running worker can read an existing queue and return a stored result. A new customer report needs another worker, and the recovery procedure needs a routing change. The current service may answer a health check while the new report and recovery path are blocked. That is a partial service result, even if every existing instance remains running.
This article proposes a dependency review for cloud and platform owners. It is not a record of an Ampity customer outage or a guarantee about a provider's failure behavior. Use the actual service, region, product version and request path in the review. Separate a provider control API failure from an application control-service outage; they can affect different dependencies and require different responses.
Locate administrative calls inside the serving path
AWS's fault-isolation guidance distinguishes administrative resource operations from a service's primary function. It gives examples such as launching EC2 instances versus using running instances, and Route 53 answering DNS queries. Our review recommendation is to inspect whether the application adds administrative calls to an otherwise established runtime path. A service boundary does not determine the dependencies of your code.
Trace a representative request from entry to accepted business result. Does the handler describe an infrastructure resource, discover a queue URL, create storage or fetch mutable configuration before doing useful work? Record the actual API, caller and failure behavior. A dependency called only at startup can affect replacement workers even when already-running workers do not use it on each request.
Inspect cached discovery results too. A stored endpoint can remove a repeated lookup, but it needs ownership and invalidation rules. How does the application learn that the resource moved or was replaced? Can an old result route a customer to the wrong environment? Removing a dependency from the hot path is useful only when the application can safely use what it retained.
Avoid classifying every configuration or identity service as the provider control plane. Some are separate runtime dependencies with their own availability and security contract. Name them individually. A diagram labelled “data plane” that omits a required authorization service can hide the operation that actually determines whether customer requests continue.
Distinguish running containers from replacement capability
Kubernetes's architecture documentation assigns scheduling and cluster-wide reconciliation to control-plane components. It describes node components maintaining running Pods and the kubelet ensuring their containers run. It also identifies controllers that populate service endpoints or interact with cloud APIs. Our operational recommendation is to test serving and replacement separately for the actual cluster setup, including its network plugins and add-ons.
A healthy container is only one part of the customer path. The request may require working ingress, service discovery, database access and external dependencies. Check those components rather than concluding that an API server outage leaves all business functions unchanged. Existing network state and a newly created endpoint can have different update requirements.
Next, ask what happens when an existing worker becomes unavailable during the outage. Can another already-running worker accept its work? Does the plan require scheduling a new Pod, provisioning a node, pulling an image or attaching storage? Record every required step and who operates it. A design that serves its usual load but cannot absorb a second failure has a narrower continuity envelope than a general high-availability claim suggests.
Do not use an outage rehearsal to justify bypassing cluster safety controls. A direct node action or manual resource edit can conflict with the intended state later. If an emergency path exists, it needs separate authorization, scope, evidence and reconciliation instructions. A workaround that makes one request succeed can create a delayed failure when controllers resume.
Write a dependency map that exposes change requirements
The map below is a proposed review worksheet. Its four rows describe questions to investigate, not universal provider behavior. Replace each evidence statement with an observation for the named application and outage boundary. If a dependency has not been tested or documented at the required scope, record it as unknown.
| Customer path | Dependency to inspect | Evidence required | | --- | --- | --- | | Serve an existing request | Runtime route, data access and current authorization | Business result and negative authorization check during the scoped outage | | Start a new session or renew access | Identity issuance, refresh and policy freshness | New-session and refresh results, with explicit expiry and revocation behavior | | Replace or expand capacity | Scheduling, infrastructure creation and startup dependencies | Replacement attempt through the approved path, with observed blocking step | | Move traffic or recover service | Routing change, writer authority and operator access | Authorized transition plus post-change business and reconciliation checks |
Include the observation time and the state that made the result possible. A successful request using an already-open connection does not establish that a new connection can be opened. A worker that already holds configuration does not establish that a restarted worker can fetch it. These are separate test cases, not inconsistent observations.
Mark dependencies that the team cannot inspect. A managed operation may hide internal components while exposing a documented API contract and observable result. Use that evidence rather than drawing an invented internal topology. The map should help the operator choose a safe action with the information available, not imply knowledge the organization does not have.
Define the time and security limits of retained state
Configuration, endpoints, certificates and credentials can remain usable for different periods. Review when each must change, who owns that change and what happens if refresh fails. Do not give them a single assumed cache lifetime. A short outage exercise may finish before the first renewal boundary and therefore leave the longer failure path untested.
Authorization needs an explicit decision too. A retained permission may help a request continue, but it may no longer reflect a user's current access. The application owner must decide which functions can use retained state and which must stop when authoritative checks are unavailable. Avoid turning an availability objective into an undocumented fail-open policy. Include a negative test with a disallowed operation, not just successful requests by an authorized user.
Describe the limitation to the customer where it affects the business function. A service might permit reading existing reports while delaying creation of new reports or account changes. The interface should name that state and preserve the user's intended request without implying that it was accepted for execution. If delayed work can expire, record how the user learns its final disposition.
AI-assisted workflows need the same boundary. An agent may propose a cloud configuration change, but the proposal does not create infrastructure or prove that the control API accepted it. Bind any permitted action to a stable operation identity and approved authority. Preserve an uncertain result for reconciliation instead of having the agent repeatedly submit changes until a response appears successful.
Rehearse failure and return without changing production
Use an isolated environment and an approved, narrowly scoped fault mechanism. Identify which endpoint or dependency the test will make unavailable and verify that unrelated production identities, networks and services are outside that scope. The exercise should distinguish a control-service failure from a total network failure. If the mechanism blocks both administrative and serving traffic, it cannot prove the intended separation.
Define stop conditions before the fault. Include maximum duration, workload limits, dependency load and a tested way to remove the fault. Assign an operator who can restore the environment independently of the blocked interface. Do not begin an exercise whose only reversal method depends on the component being deliberately made unavailable.
Test the four customer paths with synthetic identities and data. Observe established requests, new access, a bounded replacement attempt and the separately authorized recovery path. Exercise refresh boundaries through an approved test method rather than copying production credentials or changing their lifetime. Preserve denied and uncertain results alongside successes; do not turn an unattempted operation into a green status.
After access returns, inspect pending changes and delayed reconciliation. Did a timed-out request complete later? Did controllers apply a change made before the outage? Is the worker fleet, routing and application state what the service owner expects? Resume consequential automation only after the named owner checks that record. Blindly replaying every failed-looking control request can create duplicate resources or reverse a manual recovery action.
Accept a bounded continuity claim
The review output should name the customer function, permitted load, fault boundary, retained state, tested duration, blocked changes and recovery owner. Attach the dependency map and observations. A statement such as “existing report reads continued under this tested condition” can be defensible. “The platform works without its control plane” usually hides more scope than the evidence covers.
Decide what needs to change using the observed failure. An administrative lookup in each request may need removal from the hot path. An untested startup dependency may need an approved readiness exercise. Replacement capacity may need a different provisioning strategy or a narrower accepted outage mode. Compare the cost, freshness and security trade-offs before choosing more pre-provisioned resources as a universal solution.
The next useful step is one dependency map for a customer function that matters during an incident. Bring its unknown refresh and replacement paths to a reliability review. Use the DNS and client-recovery playbook for the traffic transition boundary, and the regional unsettled-write article when recovery can leave external business effects unresolved. These resources can be read without submitting an enquiry.