A Kubernetes Disruption Budget Does Not Create Spare Capacity
Check eviction permission, replacement placement and application recovery separately. Work through a node drain blocked by required pod anti-affinity.
An allowed eviction is not a replacement plan
A cluster has three healthy application replicas and a disruption budget requiring two to remain available. An operator starts maintenance on one node. The eviction succeeds, but the replacement pod stays Pending. The budget is doing its job: it has limited the permitted disruption. It has not supplied another place for the pod to run.
Before maintenance, review three different conditions. The eviction must be permitted through the chosen maintenance path. The replacement must have an eligible placement with the resources it requires. The application must become ready and serve its required workload after the move. A positive result for one condition cannot establish the others.
This article uses a fictional stateless Deployment to explain that distinction. Node names, requests and observations are synthetic, not an Ampity customer incident or an executed cluster exercise. The example is intentionally constrained so the scheduling problem can be inspected. Real maintenance decisions require the actual manifests, scheduler configuration, cluster version, workload behavior and hosting provider's drain procedure.
Know which disruption the budget actually controls
Kubernetes's disruption documentation explains that PDB-aware tools use the Eviction API. It also distinguishes involuntary losses, which PDBs cannot prevent, and notes that direct pod deletion bypasses a budget. Workload rolling updates have their own controls rather than being limited by PDBs. Our operational recommendation is to inspect the exact maintenance mechanism, not assume that every operation called a disruption obeys the same policy.
For the example, there is one policy/v1 PodDisruptionBudget, matching exactly the labels of one Deployment in one namespace. Desired replicas are three, and all three pods are Ready. The PDB uses an integer minAvailable of two. There are no overlapping budgets, outstanding disruptions or concurrent workload updates. These assumptions matter when interpreting the initial status.
The illustrative budget is small, but it is not a manifest to apply blindly:
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: analysis-api-budget
namespace: maintenance-demo
spec:
minAvailable: 2
selector:
matchLabels:
app: analysis-apiKubernetes's PDB configuration guide describes the selector, integer and percentage settings, and status fields including currentHealthy, desiredHealthy and disruptionsAllowed. It counts healthy pods through their Ready condition. Our example starts at three healthy pods and two desired healthy pods, with one permitted disruption after the controller has reconciled this state. This is not a universal arithmetic replacement for reading current status.
Verify that the observed generation corresponds to the intended policy and that the selector matches the right pods. A status from before a policy change is not current permission. A budget with the wrong selector can protect unrelated work or fail to cover the intended workload. For unhealthy pods, separately review the configured unhealthy eviction policy and supported cluster behavior; this healthy-pod example does not prescribe that policy.
Trace the replacement through every hard placement rule
The three replicas run on nodes A, B and C. Each node has a distinct kubernetes.io/hostname label and node-pool=analysis, and all application pods carry app=analysis-api. The Deployment requires nodeSelector node-pool=analysis and pod anti-affinity for matching pods across that hostname topology key, within its own namespace. Each replica must therefore be placed on an eligible pool hostname not already occupied by another matching pod. A new node with node-pool=general would fail the selector even if it had spare resources.
Node A is cordoned for maintenance, and its application pod is evicted through the PDB-aware path. Nodes B and C remain healthy, but each already hosts a matching pod. The replacement cannot use A because it is unschedulable, and it cannot use B or C because the required anti-affinity rule excludes those placements. There is no node D in this scenario.
Kubernetes's placement documentation distinguishes required rules from preferences: the scheduler filters for required affinity and anti-affinity before scoring eligible candidates. It also warns about consistent topology labels. Our engineering inference is that unused resources on an excluded node do not provide placement capacity for this replacement.
Suppose each pod's total effective request is 500 millicpu and 512 MiB after accounting for all applicable container requests and overhead. B and C each have 1,500 millicpu and 3,584 MiB of unallocated request capacity. Either has enough requested CPU and memory for the replacement, but neither satisfies anti-affinity. This is a placement failure, not evidence that the application needs lower CPU requests.
Kubernetes's resource-management documentation explains that scheduling uses requested resources; low observed utilization does not make an otherwise failing capacity check pass. Our operational recommendation is to inspect effective requests against eligible node allocatable resources, not compare a pod request with a dashboard's average usage. Account for other workloads and applicable overhead rather than treating the machine's advertised size as available pod capacity.
Other applications add further restrictions: node selectors, required node affinity, untolerated taints, storage placement or a device requirement. Do not call a node compatible merely because one check passes. For an AI model-serving workload, a spare general-purpose node may not satisfy the requested accelerator and device configuration. Adding replicas or increasing a general worker pool can leave the original bottleneck unchanged.
This example does not apply unchanged to a StatefulSet or a quorum-based service. A replacement may also need its persistent volume attached in an eligible location, application membership restored and data consistency checked. A vacant compatible hostname cannot prove those conditions. Extend the acceptance record around the actual workload rather than carrying over the stateless example's recovery conclusion.
Stop the maintenance sequence at the missing evidence
After the first eviction has settled, B and C provide the two Ready replicas. The replacement remains Pending, and the reconciled budget allows no additional healthy-pod disruption under the stated assumptions. A second PDB-aware eviction is blocked. That block is useful evidence that the maintenance plan cannot safely continue as originally intended; it is not a reason to bypass the budget.
Do not weaken anti-affinity or lower minAvailable simply to make the drain finish. Either change alters the workload's failure exposure. A permitted alternative needs the application owner's decision, a reviewed change path and evidence that the altered configuration can carry the required load. Making three pods share two hosts can make a maintenance dashboard green while reducing resilience to the next host failure.
One possible remedy is a fourth Ready, schedulable node with correct labels, compatible configuration and sufficient allocatable resources. It must satisfy every required placement condition, not just exist in a cloud console. The replacement can then be scheduled, start and become Ready. Only after observing the recovered workload and current budget should the operator decide whether to continue the next maintenance step.
An autoscaler requesting that node is not evidence of usable capacity. Provisioning can fail, startup can be delayed, and a new node can lack the required labels, devices or storage access. Record the actual blocking step. If the maintenance window cannot accommodate recovery, stop through the approved procedure and preserve the current serving replicas rather than escalating disruption merely to meet a schedule.
The five cases below are independently stated expectations for a rehearsal. They are not observed output from a Kubernetes cluster. Use them to distinguish policy status, scheduling evidence and application acceptance in the review record.
| Observed condition | Maintenance decision | Evidence still required | | --- | --- | --- | | Three Ready replicas; budget permits one eviction | First eviction is policy-eligible | Replacement placement and application recovery remain unproven | | A cordoned; matching pods occupy B and C | Stop further disruption while replacement is Pending | Another node satisfying required anti-affinity and resource checks | | D exists but has incompatible placement labels | Do not treat D as replacement capacity | Correct approved node configuration and eligible scheduling result | | Replacement on D is Running but not Ready | Do not claim restored healthy capacity | Ready condition and representative application result | | Replacement is Ready and budget permits another eviction | Consider the next separately authorized step | Current status, serving capacity and no competing maintenance |
Check what the remaining replicas can actually serve
Two Ready replicas do not necessarily sustain the current customer workload. Ready is the condition used by this budget, not an end-to-end service-level proof. Review what the readiness check measures, how traffic reaches the remaining replicas and whether the business path still has its required dependencies. An endpoint that returns a simple health response may miss a broken database operation or overloaded inference backend.
Measure representative requests during the reduced-replica period. For an analysis service, the owner may require authorized requests to produce accepted results within a defined time while disallowed requests remain denied. Preserve queue age, error rates and downstream pressure alongside replica state. A fall in throughput can precede a probe failure, and a growing queue can hide behind an otherwise healthy deployment.
Maintenance can also interrupt in-flight work. Define the application's shutdown and handover contract before relying on graceful termination. A pod that disappears cleanly from Kubernetes may have accepted a consequential tool request whose destination outcome remains uncertain. Keep the business operation identity and reconcile through authoritative destination evidence rather than assuming that an evicted worker performed no write.
An AI operations assistant can help summarize pod events, effective constraints and approved options. It must keep observed facts separate from inferred causes, and it should not autonomously delete pods, relax safety controls or widen its own permissions to make progress. A suggestion to add D is not an added node, an eligible placement or accepted recovery evidence.
Rehearse one blocked placement without touching production
Use an isolated cluster, synthetic identities and a representative test workload. Record the version, scheduler settings, exact manifests and node labels. Establish the owner, maximum disruption, abort condition and independent recovery path before exercising a drain. This article's YAML is only the budget portion of an example; it is not a complete rehearsal environment.
Have reviewers define the five expected outcomes before observing the run. Reproduce the three-host required anti-affinity arrangement, remove one node from eligibility through the approved test path and inspect the pending replacement's events. Then introduce a separately approved compatible node and observe scheduling, startup, Ready state and an application request. Do not replace these observations with a controller's desired replica count.
Include deliberately broken interpretations in the review: treating disruptionsAllowed as spare capacity, treating unused CPU as a valid anti-affinity placement, treating a provisioned node as Ready, and treating Running as recovered service. The evaluator should reject those interpretations. A rehearsal that cannot distinguish them is not proving the boundary the maintenance team needs.
These are proposed cluster tests, not executed maintenance or scheduler validation. Local article checks can validate wording, fixture parity and metadata, but cannot establish actual eviction behavior or application continuity. Separate a drained node, a stopped maintenance sequence and restored service in the final evidence. Retain unknowns rather than forcing every field into a success or failure label.
Make maintenance readiness a shared acceptance record
Before scheduling the next maintenance window, collect the intended disruption path, matched PDB and current status, owning workload, replacement constraints, eligible capacity, startup dependencies and representative serving result. Include observation times and concurrent operations. A stale successful drain from a smaller workload is not proof for today's manifests or load.
Agree who may pause maintenance and who may approve a configuration change. The cluster administrator can inspect infrastructure, while the application owner must define acceptable reduced capacity and correctness. Neither role should have to guess what a green replica count means commercially. A concise acceptance record makes that decision inspectable without promising that Kubernetes prevents every outage.
Read the control-plane continuity guide for replacement dependencies during an administrative failure, and the queue admission guide for work that can expire while capacity is reduced. A Kubernetes reliability review or cloud reliability assessment can connect maintenance controls to a defined application outcome.
These resources are available without providing an email address. If you want Ampity to contact you, use the optional request and describe the workload, maintenance change and unresolved placement constraint. Do not submit kubeconfig files, tokens or private cluster details through a public form. Begin with evidence about one replacement path before expanding the claim to general cluster availability.