Rehearse an ALB Zonal Shift Without Mistaking Routing for Recovery

Prepare a bounded ARC zonal-shift rehearsal for one Application Load Balancer. Separate applied shift state, connection cohorts, remaining-zone capacity and verified...

Scope: a routing rehearsal, not an Availability Zone outage

An ARC zonal shift is a temporary traffic mitigation for a supported resource within one AWS Region. It does not manufacture spare capacity, restore corrupted state or establish that an entire application survived an AZ outage. Rehearse the exact resource and user journey, then accept only the behavior demonstrated by the evidence.

This playbook specifies an unexecuted educational rehearsal for one already-enabled sandbox Application Load Balancer (ALB). It assumes a reviewed multi-AZ target group with cross-zone load balancing enabled. No AWS resource, shift, traffic generator or customer application was tested for this reference. Vendor behavior was rechecked October 8, 2026. Production use requires a separate change authorization and workload-specific review.

The broader Multi-Region vs Multi-AZ playbook owns topology and data-authority choices. Here, the operational task is narrower: bind one manual shift to applied-state readback, measured traffic behavior, remaining-AZ outcomes and a safe return. Do not promote a database writer or move Regions as part of this exercise.

1. Bind the rehearsal to one resource and operator

Owner: service owner and load-balancer operator. Output: signed-off rehearsal manifest. Record account, Region, ALB ARN, enabled zones and their identifiers, target-group ARNs, listener rules and deployment revision. Include the synthetic journey, expected result, offered work, excluded paths, collection window and maximum cost. Use references to credentials, never credentials themselves.

The security reviewer separates read-only inspection from mutation authority. Hand off the ARC zonal-shift authorization reference with this procedure's three actions: arc-zonal-shift:GetManagedResource for readback, arc-zonal-shift:StartZonalShift for start and arc-zonal-shift:CancelZonalShift for cancellation. These support ALB resource scope. Record the exact requester, reviewed ALB ARN, relevant conditions and effective-permission review evidence in the manifest. Do not infer scope from the cancellation command's shift-ID input alone.

Separately list the load-balancer inspection and telemetry access required by the chosen evidence plan, such as approved attribute, target-health and metric/log reads. Their authorization is not supplied by the three ARC actions. The security owner resolves those permissions and applicable restrictions before start. This is a scoped handoff, not a universal IAM policy or permission grant. A policy review or simulation is not an observed successful start or cancellation. Do not fix a permission error by requesting unrestricted administrator access. A second operator must know the shift identifier, cancellation route and escalation path.

Gate: if identity, cancellation access or the resource boundary cannot be verified, stop before changing traffic. An operator who can start a shift but cannot cancel it is not a ready rehearsal owner. Record why the gate failed and what evidence would clear it.

2. Check ALB eligibility and exclude misleading topologies

Owner: network/load-balancer operator. Output: applicability row. Inspect current ALB and target-group attributes, registered targets and zonal placement. Existing opt-in is a prerequisite, not an implied instruction to enable it during collection. Current ARC ALB guidance supports cross-zone-enabled and cross-zone-disabled configurations, but they must not be treated as one experiment.

For this cross-zone-enabled fixture, AWS describes removal of the zonal ALB address from DNS and blocking traffic to targets in the shifted AZ. A single-AZ target group is not a valid shift target for this procedure. If an NLB fronts the ALB, the documented shift belongs at the NLB; reject this ALB-only rehearsal rather than assuming the upstream NLB recognizes it.

Gate: the operator verifies the exact supported resource and Region, all relevant target groups and routing rules. Mixed listener paths or a single-zone dependency are named limitations, not hidden behind a healthy default target group. EKS, Auto Scaling, NLB and cross-zone-disabled rehearsals need their own behavior review and are outside this article's fixture.

3. Inspect other shifts before choosing the away-from zone

Owner: authorized operator. Output: timestamped managed-resource snapshot. Use the managed-resource read API for the exact ARN and Region. The sample below is read-only, but its placeholder values must be replaced and independently checked before an authorized operator uses it.

aws arc-zonal-shift get-managed-resource \
  --resource-identifier 'REVIEWED_SANDBOX_ALB_ARN' \
  --region 'REVIEWED_REGION' \
  --profile 'APPROVED_READ_PROFILE' \
  --output json --no-cli-pager

Keep the full permitted response. Review active zonal shifts, autoshifts and applied weights rather than selecting the first list item. AWS distinguishes an active shift from the shift currently applied to a resource. Customer shifts take precedence over autoshifts, which take precedence over practice runs. An ACTIVE response alone therefore cannot establish which traffic treatment is in effect.

Use the linked response schema as a field checklist: match top-level arn to the reviewed resource; retain each relevant zonalShifts[] entry's zonalShiftId, resourceIdentifier, awayFrom, expiryTime and appliedStatus. Inspect autoshifts[].appliedStatus and awayFrom separately. Retain the entire appliedWeights map with its zone keys. These weights describe zonal activity, not measured request share or correct user completion. Classify the intended treatment as APPLIED, NOT_APPLIED or unknown/ambiguous; only a matching applied action clears this treatment gate. A missing entry is not an invented NOT_APPLIED response.

Gate: if another shift, practice run or an unexplained absent zone affects this sandbox, pause and resolve ownership. Do not cancel another team's action to make the fixture tidy. Record the selected zone using the actual mapping for this account; a remembered zone label from another account is not a checked identity.

4. Prove the baseline before withdrawing capacity

Owner: application owner and SRE observer. Output: baseline evidence bundle. Run the synthetic journey at the approved workload mix. Verify its known result, completion deadline and authoritative readback where it writes state. Show that all intended zones are serving the expected path and that dashboards can distinguish them.

Capture offered requests, correct completions, errors, latency distribution, retries, active connections and dependency constraints. Record metric dimensions, statistic, units and period with each export. Regional totals can conceal a weak zone. Use application instrumentation when an ALB metric cannot establish correctness or the target's placement. Do not expose private topology or customer identifiers through a new public response header.

Gate: a broken baseline, missing telemetry or an unclassifiable result stops the rehearsal. A load-balancer health check is not the known business result. Preserve any difference between a newly established client and a long-lived process before attributing later differences to the shift.

5. Calculate remaining-AZ capacity without assuming even traffic

Owner: capacity owner with dependency owners. Output: reviewed capacity worksheet. The ARC overview requires sufficient capacity before traffic is withdrawn. In this playbook, capacity means demonstrated useful workload under the accepted latency/correctness boundary, not an instance count multiplied by a guessed rate.

Use this fictional arithmetic only to check the worksheet. Suppose three zones each have an independently established useful capacity of 120 request-units per second for the stipulated workload mix. The admitted offered load is 180. Losing one zone leaves 240 units of nominal combined capacity, with an aggregate allowance of 60. This is not a benchmark, recommended operating margin or evidence that real traffic divides equally.

If one remaining zone receives 140 and the other receives 40, the first exceeds its stipulated 120 even though 180 is below 240. A shared database capped at 150 can also invalidate the application total. Record distribution, background work, connection ceilings, cache warming and dependency allowances separately.

Illustrative assumptions only:
per-zone useful capacity: 120 units/s
three-zone capacity: 360 units/s
one zone withdrawn: 240 units/s
offered work: 180 units/s
aggregate allowance: 60 units/s
skewed remaining load: 140 + 40 = 180 units/s
first remaining zone excess: 20 units/s
AWS test result: NOT EXECUTED

Gate: approve only a workload envelope supported by actual evidence for the selected configuration. If new capacity depends on unverified scaling during the event, readiness remains conditional. An operator can choose a lower approved synthetic load for an initial test, but must state that it does not demonstrate peak readiness.

6. Separate new connections from existing ones

Owner: client/application owner. Output: connection-cohort manifest. Include a fresh-connection cohort and a representative persistent-connection cohort. State which client library owns reconnects, connection pooling and retries, plus how each cohort's route and outcome are observed. Use the same reviewed workload semantics for both.

AWS's ALB attribute guidance distinguishes client keepalive duration from connection idle timeout. Existing clients can retain their initial keepalive value after an attribute change; new connections receive the new value. Do not lower a timeout during this exercise and claim all earlier clients immediately inherited it.

An ALB may still receive an existing client connection through the excluded zonal node while target routing follows its cross-zone-enabled shift behavior. Distinguish the client ingress location from the target serving the request. A DNS lookup cannot prove both. Record unresolved old connections rather than inferring full drain from a changed answer.

Gate: attribute tuning, forced disconnects and client changes require a separate reviewed candidate test. Preserve the original configuration. Do not introduce them simultaneously with the first manual shift and lose the ability to explain the result.

7. Write stop, expiry and return rules before starting

Owner: service owner, operator and observer. Output: executable stop card. Select a finite expiration and an observation window that fits inside it, leaving time to cancel and validate return. Start-zonal-shift requires resource, away-from zone, expiration and comment. The documented initial lifetime is bounded by 72 hours; that maximum is not a recommendation to run a sandbox test for three days.

Express the approved expiration as a positive whole number followed by m or h, within the current API bounds. Do not pass the literal placeholder shown below. The stop card names the authorized cancel operator, independent expiry watcher, page/escalation route and signals that end exposure. Include incorrect results, deadline violations, remaining-zone saturation, retry amplification, cost limit, loss of telemetry and wrong resource/zone readback. The team fills actual thresholds before approval; a blank threshold is not a pass.

Gate: cancellation returns traffic toward the original configuration, so the original zone must be fit to receive it. In this healthy sandbox fixture, verify that condition before starting. If a real impairment persists, blindly returning traffic can worsen the incident. Use the separately approved incident mitigation, not this demonstration's return assumption.

8. Start one approved manual shift and retain its identifier

Owner: mutation-authorized operator. Output: shift ID and start/expiry record. The following is a command shape, not a command executed by the author or a grant of authority. A reviewer must bind every placeholder to the manifest and the approved expiration. Keep the comment free of confidential customer data.

aws arc-zonal-shift start-zonal-shift \
  --resource-identifier 'REVIEWED_SANDBOX_ALB_ARN' \
  --away-from 'REVIEWED_AZ_IDENTIFIER' \
  --expires-in 'APPROVED_DURATION' \
  --comment 'APPROVED_REHEARSAL_REFERENCE' \
  --region 'REVIEWED_REGION' \
  --profile 'APPROVED_CHANGE_PROFILE' \
  --output json --no-cli-pager

Immediately retain the returned shift ID, resource, away-from value, status and expiry. Independently repeat the managed-resource readback and determine whether the intended shift is applied. If the start request's result is lost or ambiguous, inspect current state before issuing another mutation. Do not keep restarting or extending the shift to make a dashboard look cleaner.

Gate: a mismatched ARN, zone or applied treatment stops the planned observation. Escalate and follow the reviewed cancellation route for the identified action. Operator/API success is only the first evidence boundary.

9. Observe four proofs rather than one green status

Owner: SRE observer and application owner. Output: aligned timeline. Keep command acceptance, applied state, observed traffic behavior and correct user completion as separate timestamps and evidence links. Add return validation as its own stage. The reference figure is a proposed evidence progression, not traffic topology or a tested deployment.

Four separate proofs: the intended ARC shift is applied, ALB traffic behavior is observed, the remaining zones preserve correct user outcomes, and the return path is validated. Missing evidence at any proof holds acceptance for owned investigation.

Blue arrows are evidence progression, not network traffic or automatic approval. Brown branches hold acceptance when a proof is missing. The last proof is required even if reduced-zone service appeared healthy. No shift or user result was executed for this figure.

Compare the fresh and persistent cohorts, zonal target load, correct completions and shared dependencies against baseline and the stop card. A reduction in one zone's requests does not establish the remaining zones are within capacity. A rise in retries may make admitted work appear stable while useful work deteriorates.

Gate: advance only if the planned evidence is available and compatible. If it is not, report an instrumentation gap or failed condition. Do not substitute lower offered work, a wider deadline or a different journey after the test began without creating a new reviewed experiment.

10. Cancel, then verify actual traffic return

Owner: authorized operator and application observer. Output: cancel response plus return record. Use the retained, independently checked identifier. The cancel API can cancel customer-started shifts and practice-run shifts, not an arbitrary AWS-initiated autoshift. This article authorizes neither kind of real action.

aws arc-zonal-shift cancel-zonal-shift \
  --zonal-shift-id 'VERIFIED_REHEARSAL_SHIFT_ID' \
  --region 'REVIEWED_REGION' \
  --profile 'APPROVED_CHANGE_PROFILE' \
  --output json --no-cli-pager

Read the managed resource again. Another active action can still determine applied routing after this shift is canceled; do not equate one CANCELED result with full restoration. Observe intended zonal participation, both client cohorts, correct outcomes and the agreed settling window. Expiry is a temporary-action safety boundary, not a complete return test.

Gate: if return does not meet the agreed conditions, retain the incident/rehearsal owner and investigate. Cancellation cannot reverse completed writes, external effects or unreviewed configuration changes. Any rollback of a separately approved attribute candidate must restore its reviewed values and recheck old/new connection cohorts.

11. Keep autoshift admission outside the manual result

Owner: service owner. Output: separate automation decision, if requested. A successful manual rehearsal is useful input, not approval to enable autoshift. Current autoshift guidance describes AWS telemetry-based decisions rather than inspection of each individual application's health, plus required practice runs. The automation does not wait for capacity scaling to complete.

If the team later proposes autoshift, review its capacity, alarm coverage, practice-run behavior, notifications, exclusions and operating ownership separately. A practice-run outcome is scoped to its configured controls; it is not proof that an uninstrumented write, scheduled worker or business invariant remained correct.

Gate: do not change practice-run configuration or automation enrollment while trying to explain this first manual fixture. The general reliability playbook remains the owner for wider dependency exercises and accepted service commitments.

12. Close with a reusable decision packet

Acceptance criteria for this rehearsal

The incident lead records PASS, FAIL or HOLD for each item, with an evidence reference and named follow-up owner. A successful API response alone cannot pass this checklist.

  • The exact ALB, Region, away-from zone, eligible target-group topology and authorized operator match the approved packet. Conflicting shifts and precedence have an explicit disposition.
  • Baseline user checks, residual-zone capacity and stop thresholds were accepted before withdrawal. Actual offered load and the observation window are retained, not inferred from the illustrative calculation.
  • The retained shift ID and resource response establish the intended action was applied. Zone-keyed applied weights are not represented as measured request shares.
  • New-connection traffic observations and existing-connection behavior support the stated routing conclusion. Required user outcomes remain within the agreed thresholds; missing coverage is HOLD.
  • Cancellation or expiry is followed by dated applied-state, traffic and user-outcome readback. Temporary alarms, retained evidence and residual conditions have cleanup or explicit owners.

Record the overall disposition separately: accepted manual rehearsal, failed rehearsal with evidence, or inconclusive. Passing a manual rehearsal does not approve autoshift or establish an Availability Zone outage result.

Owner: service owner and receiving operator. Output: accepted tested scope or owned gaps. Use the record below. Values remain pending until observed; this article supplies no fabricated reviewer or successful customer result.

Rehearsal identity / revision / NOT EXECUTED or execution date:
Account / Region / ALB ARN / target groups / cross-zone settings:
Selected away-from zone / other shifts and applied precedence:
Owners / start-cancel authority / expiry watcher / escalation:
Requester / ALB-scoped ARC actions / conditions / permission-review evidence:
Separate ALB inspection and telemetry permissions / security owner / gaps:
Synthetic journey / expected result / offered work / excluded paths:
Remaining-zone and shared-dependency capacity evidence:
Fresh and persistent connection cohorts / initial keepalive values:
Stop thresholds / maximum spend / return readiness:
Start response / shift ID / expiry / applied-state readback:
Managed-resource arn / matching zonalShiftId / resourceIdentifier / awayFrom:
Selected expiryTime / zonalShifts[].appliedStatus / capture time:
Other zonal shifts and autoshifts[].appliedStatus / awayFrom:
appliedWeights with zone keys / APPLIED, NOT_APPLIED or unknown classification:
Traffic evidence / correct completions / latency / retries:
Stop or cancel decision / response / remaining applied shifts:
Return evidence / temporary configuration cleanup:
Accepted scope / failed or unknown proof / owner / next test:

Classify the outcome as observed success within the stated envelope, failed condition, inconclusive instrumentation or not executed. Include confounders: artificial load distribution, already-warm capacity, omitted dependencies and shorter-lived clients than production. Retain contradictory evidence rather than turning the record into a generic resilience testimonial.

Done means the intended treatment and return were observed, the chosen journey met its agreed rules in the tested envelope, and unresolved conditions have owners. It does not mean every AZ impairment is safe, peak capacity is established by low load or a customer availability promise has been met. The next action is to complete the manifest and capacity/stop card for one eligible sandbox resource before requesting a mutation window. For cross-team help, a reliability review can scope that evidence; it cannot substitute for it.

Related resources

Related services