Cross-Zone Traffic Can Hide in a Healthy Architecture

Evaluate cross-zone traffic without sacrificing availability. Compare routing, capacity and dependency placement in normal operation and zone-failure tests.

A healthy service can still take an expensive path

Cross-zone traffic is not automatically waste. Some flows preserve redundancy or let requests reach available capacity. Others cross an Availability Zone because a route, dependency or worker placement changed unnoticed. Before reducing that traffic, identify the actual path and prove that the proposed change preserves the service's failure behavior.

A green availability dashboard does not explain network placement. The application may meet its latency target while sending repeated large responses between zones. Conversely, a low-transfer design may overload one zone when local targets disappear. Cost and reliability need a shared decision record rather than separate optimizations that contradict each other.

This article proposes an AWS-oriented review. The worker example and comparison sheet are synthetic, not measured customer outcomes. It does not provide current service rates or a universal recommendation to disable cross-zone routing. Use the applicable service configuration, billing treatment and accepted availability requirements for your workload.

Separate three different reasons for crossing a zone

First, a load balancer may send a request to a target in another enabled zone. Second, an application in one zone may call a database, cache or peer service elsewhere. Third, a worker's outbound route may use infrastructure placed in another zone. Changing the first path does not necessarily change either of the others.

AWS describes cross-zone load balancing as distribution across targets in enabled zones, compared with distribution limited to the receiving node's zone when disabled. Its documentation also distinguishes configuration behavior across load-balancer types. Inspect the actual load balancer and target-group settings rather than treating “cross-zone” as one account-wide switch. See how Elastic Load Balancing routes requests.

Keep the application dependency map beside the routing configuration. A zone-local frontend can still make every database call across a zone. A local cache can reduce normal traffic but introduce cold-cache load during failover. The relevant design boundary is the complete request or job, not only its first hop.

Identify who owns each path. Platform engineers may own subnet routes while an application team controls peer discovery and payload size. A cost recommendation without those ownership distinctions can be implemented at the wrong layer, leaving the expensive flow unchanged while weakening unrelated redundancy.

Inspect effective placement, not intended placement

Inventory the producer and destination zones during the disputed time window. Record effective routes, target membership and dependency endpoints. Infrastructure code describes intended configuration; it does not by itself prove where every running task was placed or which endpoint a long-lived process used.

For shared or short-lived resources, retain the historical mapping needed to associate traffic with a workload. Today's healthy task list cannot reliably identify a worker that terminated yesterday. Avoid assigning a flow to a team solely from an address that may have been reused.

Check request direction and selected measurement boundary. One payload can appear at several observation points. Do not add all those observations and call the sum chargeable transfer. Resolve the billing category separately, then choose observations that help test the suspected path. Document the remaining uncertainty rather than forcing an exact match.

Also distinguish payload growth from path growth. A larger response, extra retries and a new cross-zone route can all raise observed traffic. Compare completed business outputs, attempt counts and payload sizes before deciding that placement alone caused the increase.

Work through a worker fleet with shared outbound routing

Suppose a synthetic batch service runs workers in two zones but uses one zonal NAT gateway. Both worker groups successfully call an external partner, so the service appears healthy. A new export workload increases traffic from workers outside the gateway's zone, making the existing placement decision more consequential.

AWS's NAT pricing guidance identifies hourly availability and data-processing charges, and suggests same-zone placement or a gateway in each resource zone when significant traffic crosses zones. It also suggests considering supported endpoints for suitable AWS-service traffic. These are options to evaluate, not proof that adding infrastructure always lowers the total bill. See AWS's NAT gateway cost guidance.

Compare retaining the current path, adding zone-local outbound paths and changing eligible service access. Include fixed infrastructure costs and operating complexity, not just the transfer component expected to decrease. The external-partner flow cannot be assumed to use an AWS-service endpoint merely because another part of the workload can.

Test the partner's access requirements too. An outbound-path change can alter the source address seen by the partner. If an allowlist or identity check depends on that address, the proposed optimization may break delivery despite healthy internal workers. Record and verify those dependencies before switching routes.

Compare normal operation with zone impairment

Use the same business-output definition across the options. For a batch service, that might be a successfully delivered export rather than a completed worker invocation. A retrying worker can look active while delivering nothing. Normalize observations against useful outputs and retain failed or delayed work in the comparison.

| Review condition | Cost evidence | Reliability evidence | | --- | --- | --- | | Normal matched workload | Charged categories and observed path | Delivery completeness and latency | | Uneven target capacity | Redistribution and repeated attempts | Per-zone load and queue growth | | One zone unavailable | Changed traffic and remaining resources | Accepted outputs and dependency reachability | | Returning zone | Rebalancing and cache warm-up | Recovery stability and duplicate-work checks | | New route rejected by partner | Failed attempts and retries | Clear failure status and reversal path |

The table is a review sheet, not a benchmark result. Fill it with evidence from your own controlled tests. Name the observation window, configuration revision and workload used, so a later reviewer can distinguish a comparable test from an unrelated quiet period.

Do not evaluate zone-local routing only when both zones have abundant capacity. Include an uneven target distribution and a loss of local capacity. The surviving path must have enough capacity and reachable dependencies for the service's accepted failure mode. A cheaper normal-state route is not acceptable if it violates that operating requirement.

Use non-production rehearsal where feasible and define safety limits for any production exercise. Do not remove a zone from a live workload merely to prove an article's hypothesis. The incident procedure, authorization and customer consequences belong to the service owner.

Make locality a preference only where the contract allows it

For some application calls, preferring a local healthy peer can reduce unnecessary crossings while retaining a tested remote path. But locality must not override consistency, authorization or the location of the authoritative writer. A nearby replica is not automatically valid for a read that requires current committed state.

If the application introduces zone-aware selection, define what happens when local capacity is unavailable or stale. Bound retries and avoid concentrating every fallback onto a single surviving target. Observe queue length and completed outputs, not just connection success, during the transition.

Local caching introduces another trade-off. It may reduce repeated traffic while increasing warm-up work and the risk of serving stale material. Identify the cache's validity contract and the cold-start dependency before attributing every byte avoided to a successful optimization.

This recommendation does not apply uniformly to all managed services. Their routing, replication and charge treatment differ. Inspect the supported controls and service-specific documentation. A generic instruction to “keep everything in one zone” discards the reliability purpose of multi-zone placement and is not a responsible architecture rule.

Define acceptance and reversal before changing the path

Write the decision as an operating contract: the change should reduce a supported cost driver while preserving specified delivery and failure behavior. State the workload assumptions, expected traffic path, fixed-cost additions and unresolved allocation. Do not claim savings before the relevant billed usage has been observed.

Define reversal triggers such as delivery failures, partner rejection, sustained queue growth or inability to use the recovery path. Record the original configuration and ensure the rollback procedure does not abandon pending jobs. A route can be restored while some workers retain old connections, so verify recovery with the actual client population.

Compare matched follow-up windows and explain demand changes. A drop in transfer during lower demand is not proof of a better path. Equally, increased total traffic may coexist with lower traffic per completed output if the service is doing more useful work. Keep both views available.

Treat findings as workload-specific. A zone-local route accepted for one batch service is not automatic approval for every transactional application. Reuse the review method, not the conclusion, when a different workload has different consistency or recovery requirements.

Start with one crossing you can explain

The next action is to select one high-volume flow and document its producer, destination, intermediates and purpose. Establish which charge category it plausibly affects, then compare one reversible change under normal load, uneven capacity and zone impairment. Keep unsupported allocation visible throughout the decision.

For the billing investigation itself, read which data flow caused the cloud egress bill. For client behavior after routing changes, see why DNS failover is not instant client recovery. These are complementary checks, not substitutes for the workload's failure rehearsal.

Ampity's cloud reliability review can connect routing choices to business acceptance criteria. You can use this comparison sheet independently, without submitting contact details or uploading private infrastructure records to the website.