Test a Cross-Cloud Dependency Before Moving Application Traffic

Run a bounded cross-cloud latency probe with an approved call mix, request and attempt measurements, retry limits, transfer accounting and recovery gates.

Test the complete application request before moving traffic across a cloud boundary. A low ping time does not establish that the application can authenticate, fetch the right revision, finish its fan-out, retry safely and recover within its deadline. Freeze those requirements first, then run a bounded comparison whose failures are retained alongside its successes.

This playbook helps a service owner commission and review that probe. It covers a dependency called from an AWS-hosted application toward an approved external environment, with an equivalent baseline route. The procedure also applies in the opposite direction when the owners supply the corresponding permissions, meters and controls. It does not recommend a universal network topology, authorize production fault injection or establish legal permission to move data.

The example and companion are entirely synthetic and offline. No AWS API, provider endpoint, customer traffic, account data, packet capture or paid model call was used. Their results demonstrate counting and decision logic, not the speed, cost or reliability of a real cloud route. Reading and using the material requires no email. A later live probe needs its own change, data and spending authorization.

1. Commission one operating question

The service owner names the dependency and the decision the probe should inform. A useful question is whether moving a read-only quote dependency to another environment preserves the accepted quote deadline at the required demand, including cold connections and a dependency interruption. “Is cloud A faster?” leaves the route, operation and service contract undefined.

Assign an operator who can stop admission, a dependency owner who can inspect its receiving boundary, a network engineer who can resolve the actual path, and an observer who can challenge missing records. The cost owner supplies applicable rates later. The data owner approves payload fields, retention and destinations. Store a sanitized contract reference in the shared packet and restricted details in the organization's controlled evidence store.

Keep this task separate from the data-placement decision, which compares remote reads, source-local compute and a replica. The egress investigation playbook owns historical bill attribution. Here the output is a prospective request and failure envelope for one candidate path. A successful probe cannot explain an earlier bill or authorize the whole migration.

Gate: the sponsor accepts a specific question, maximum scope and possible inconclusive outcome. Hold if a favorable latency number is already promised, the receiving owner is absent or the team cannot identify who can stop the work.

2. Approve the call mix and representative fixtures

The application owner records operation classes from permitted workload evidence. Include proportions, payload-size ranges, response-size ranges, cache states, connection reuse, source revisions, fan-out and synchronization. Separate interactive lookup, composed response and large export. Their deadlines and dependency counts may differ. A small median lookup cannot stand in for an export that must return a large object.

Use inert synthetic fixtures first. Remove credentials, customer identifiers, protected documents and business destinations from the comparison environment. If later approved production-derived fixtures are necessary, record who authorized them, the allowed fields and their disposal condition. Hashes identify fixture bytes; they do not prove representativeness or permission.

Keep rare large or failure-prone classes visible even when their share is small. If the production mix is unknown, label the mixture as a scenario and avoid a production-weighted conclusion. Fault-enriched cases serve a different purpose from representative demand. Do not merge an outage drill into the ordinary mix to produce one appealing headline percentile.

Output: a frozen class register with its evidence source or explicit synthetic status. Each class has an outcome check, deadline, weight and permitted side effects. Stop when candidate and baseline use different source revisions or a class quietly disappears from the candidate run.

3. Draw the actual route and measurement boundaries

The network engineer records both endpoints, deployment locations, DNS resolver, TLS termination, authentication, proxies, gateways and any private connection. Record which hops remain unresolved. “Private” does not establish a latency guarantee, and an AWS account boundary does not identify the external provider's charging rules.

An approved fixture driver reaches an application in an AWS test boundary. Dependency attempts cross a resolved network path to an isolated external test adapter. Request elapsed time and attempt time are separate observations; the stop owner can close admission without depending on the external adapter.

Conceptual test boundaries, not a deployed AWS architecture. The response travels back over the same logical relationship, but its actual physical path and meters must be resolved separately. Only inert destinations are approved in this example.

Measure complete request elapsed time with a monotonic timer at the caller. Add attempt spans at the application boundary, retaining connection, authentication, queue, transfer and downstream intervals where instrumentation supports them. Distributed clocks may not be synchronized; subtracting timestamps from different hosts can fabricate precision. Preserve missing intervals and clock uncertainty rather than forcing the spans to sum perfectly.

The observation plan should distinguish attempted bytes, returned bytes and provider-billed units. Capture approved metadata, not whole payloads. A logical response can traverse more than one charged intermediary. The topology supplies a hypothesis; actual instrumentation and usage records must establish the observed route and charges.

4. Fix deadlines, retries and concurrency before running

The service owner gives each class a complete-request deadline and an accepted outcome definition. The operator records connection timeout, attempt timeout, maximum total attempts, retryable conditions, backoff, cancellation behavior and the layer that owns retries. Include SDK version and effective runtime settings, not only a configuration file that may be overridden.

AWS retry guidance recommends bounded retries with backoff and jitter and warns about compounding retries across layers. Treat the library and application together: an outer three-attempt loop around an inner three-attempt loop can make nine dependency attempts. That is an upper bound for those loops, not an assertion that every failure produces nine calls.

The current AWS SDK reference distinguishes total attempts from retries and describes an opt-in 2026 retry behavior. Pin the SDK version and applicable behavior before copying defaults. The example's 70 ms retry wait is an invented observation, not an AWS default or a jitter recommendation.

Define arrival rate, maximum in-flight requests, maximum in-flight dependency attempts and queue bound separately. A closed-loop generator waits for responses and may reduce offered load exactly when the path slows. An open arrival schedule can expose queueing, but must cap pending work and record rejected or missed admissions. Do not remove those records from the success denominator.

Gate: all nested retries and timeouts fit the class deadline or terminate through an accepted failure path. Concurrency limits are safety bounds, not proof of tested throughput. Stop if the receiving service cannot approve the proposed pressure.

5. Rehearse isolation, instrumentation and stopping

The operator sends a small authorized fixture cohort into the isolated adapter and verifies correlation at both ends. Check that every admitted request has a terminal or explicitly outstanding record, and every attempt belongs to a request. Exercise denied production destinations without issuing a real business transaction. Use the shadow-traffic boundary guidance when mirrors, workers or queues could regain production authority.

The observer rehearses the stop control before increasing load. Stopping admission must remain available when the dependency is unavailable. Record the last admitted request, in-flight attempts, queued work, delayed retries and any permitted diagnostic objects. A stopped generator does not establish that accepted remote work stopped or that earlier writes were undone.

Set owned thresholds for queue depth, sustained error rate, bytes, duration and spend, with the observation interval and missing-data behavior. Real billing can arrive too late to enforce a small test budget; use an approved conservative admission/byte limit as well as financial observation. If telemetry vanishes, stop or hold under the declared rule instead of calling silence healthy.

AWS FIS stop conditions use specified CloudWatch alarms to stop an FIS experiment. That is relevant only if a separately authorized experiment actually uses supported FIS actions and configured alarms. It does not install a stop control for this provider-neutral offline worksheet or undo external side effects.

6. Compare steady demand without hiding waiting work

The operator runs baseline and candidate with the same fixture revision, declared schedule and outcome evaluator. Record warm-up separately, then include cold connection, cache miss, authentication renewal and normal connection reuse where those states belong to the intended use. Interleave or repeat comparison windows when shared environmental changes could otherwise dominate the result, keeping the schedule predeclared.

Capture planned arrivals, admitted requests, rejected admission, queue time, terminal elapsed time, accepted outcomes, failed outcomes, attempts and retry waits. Drain bounded outstanding work after admission ends and account for requests that cannot reach a terminal observation. Missing requests remain unknown, not a zero-latency success.

Report a distribution for each class and tested concurrency band. Also report the mixture, but never average class p95 values to obtain mixture p95. Compute it from the actual combined observations under the declared weighting. Include success-only latency and the distribution of all terminal outcomes, with timeout counts. A path that immediately rejects everything can have excellent terminal latency and no useful service.

CloudWatch's statistics reference distinguishes percentiles from averages and documents limits when only statistic sets are published. Keep the source samples or supported distribution representation required by the selected method. Our offline fixture uses the nearest-rank order statistic, explicitly not a promise to reproduce CloudWatch's interpolation or aggregation.

7. Trace one request through retry and deadline

The observer inspects several correlated requests, including a retry, a terminal timeout and an outcome that arrives after its deadline. A dependency returning a response is insufficient if the user request timed out first. For a write-capable operation, a timeout may leave the receiving result unknown. Reconcile the operation identity before retrying or routing to a second destination.

Synthetic retry request: attempt one times out at 250 milliseconds, a 70 millisecond wait precedes attempt two, and its 180 millisecond response completes the request at 500 milliseconds. Two attempt records belong to one request; the separate request deadline is 600 milliseconds.

The invented 500 ms request includes 250 ms, 70 ms and 180 ms. The 600 ms deadline is a proposed fixture gate. Real instrumentation must include omitted queue, connection or local intervals within its timing contract.

The Amazon Builders' Library explains why timeout coverage, retry safety and network delay require care. Inspect whether the timer covers DNS, TLS and the intended response completion. Streaming first-byte latency and full-output completion answer different questions. For an AI workflow, measure accepted retrieval and complete usable output separately from first token, recording tool calls and downstream effects within the operation boundary.

Gate: attempt identities, retry waits and request outcomes reconcile. An unexplained timeout followed by a receiving-side action remains an unresolved correctness issue even if the percentile meets the numerical gate.

8. Recalculate the offline example

The companion's invented healthy cohort contains 100 requests: 60 lookups, 30 composed reads and 10 exports. Each request makes one dependency call, except four lookups that make two attempts. Fifty-six lookups finish in 120 ms. Four take 250 ms for a timed-out first attempt, wait 70 ms and finish the second attempt in 180 ms, producing 500 ms request durations. Composed reads take 300 ms; exports take 900 ms.

There are 104 attempts: 100 successful attempts and four timeouts. All 100 requests produce the required synthetic output. The contract allows lookup/composed deadlines of 600 ms and export deadlines of 1,200 ms. The overall nearest-rank request p50 is 120 ms, p95 is 900 ms and p99 is 900 ms. Class p95 values are 500 ms, 300 ms and 900 ms respectively. An average of those class percentiles would not describe the mixture.

The illustrative partial-failure cohort replaces all ten export responses with terminal 600 ms timeouts and suppresses further attempts under its predeclared deadline policy. It still contains 100 requests and 104 attempts, but only 90 accepted outputs. All-terminal request p95 is 600 ms, while success-only p95 is 300 ms. That apparent improvement cannot satisfy the export outcome gate. Preserve the failed class instead of celebrating a smaller percentile.

The outage fixture contains 100 first attempts that terminate at 250 ms with no accepted output and no retry. Its 250 ms p95 is also not a passing service result. Cost per accepted output is undefined when the denominator is zero. These cohorts are constructed arithmetic examples, not failure probabilities, measured availability or confidence bounds. Their predeclared failure handling differs, so a smaller terminal latency cannot establish a causal route improvement. A hundred invented observations cannot establish a real p99 service objective.

9. Keep transfer and useful output in the cost ledger

The cost owner associates directed bytes with attempt records. In the healthy fixture every attempt sends 2,000 decimal bytes, while only successful attempts return 8,000 bytes. That produces 208,000 outbound bytes and 800,000 inbound bytes. Four timed-out attempts still send data. Responses are excluded only by the fixture's explicit zero-return-byte assumption; a real timeout may receive partial data and must retain it.

For an invented linear projection to ten million requests with the same mix and retry proportion, the totals are 20.8 decimal GB outbound and 80 GB inbound. Synthetic rates of 0.08 and 0.02 currency units per GB produce 1.664 and 1.600 units, or 3.264 units of modeled transfer. These are not AWS or another provider's rates. Replace both directions with applicable product, location, tier, agreement and meter evidence.

The example adds distinct assumed costs: fixed network 20, intermediary processing 3, requests 2, compute 12 and observations 1 unit, totaling 38 units before transfer. The total is 41.264 units for the projection. No category is claimed to be eliminated, and no charge is inferred to be free. The fictional projection assumes unchanged routing, demand mix, retry behavior and linear costs, omitting capacity steps, commitments, taxes and FX.

AWS transfer-modeling guidance supports resolving source, destination, volume and workload benefit. Keep bytes, accepted outputs and cost basis distinct. The offline fixture cannot reconcile a provider invoice, and network observations alone do not establish every billable charge.

10. Exercise partial failure and recovery separately

The operator and receiving owner choose a bounded fault matrix before the run. Cover a slow subset, dependency unavailability, DNS or connection failure, throttling, interrupted large responses, credential failure and delayed completion after the caller stops waiting, where each applies. Change one fault condition at a time when attribution matters. Keep permission failures separate from transient failures that may legitimately recover.

Specify expected user behavior for every case: accepted cached data with its freshness label, a controlled unavailable response, a deferred job or a denied operation. A fallback must independently satisfy processing permission, source revision and outcome requirements. Switching to another AI provider or a different replica can change those boundaries. Do not enable it merely because the normal path timed out.

AWS resilience-testing guidance recommends bounded hypotheses, guardrails and restoring known-good state. Start with isolated adapters and pre-production controls. This playbook does not provide executable production fault commands. A separate approved plan must name targets, actions, blast radius and restoration evidence before any live fault exercise.

Stop: unexpected destinations, uncontrolled amplification, unmet admission limits, absent observations or unintended effects. Preserve the failed cohort and configuration. Do not retry the whole experiment until a fortunate run appears without retaining the earlier attempts.

11. Restore the baseline and reconcile outstanding work

The operator stops new admission, records the final request identity and inspects active connections, queued work and delayed retries. Restore the exact baseline route and effective configuration through the organization's change procedure. The observer runs the declared baseline fixture, confirming accepted output and timing within the tested envelope. Removing the fault is not proof that the caller has resumed healthy behavior.

Check receiving-side outcomes for requests whose completion was unknown. A routing reversal cannot retract a notification or undo a write. Assign reconciliation to the destination's owner and retain unresolved identities in the packet. Do not delete shared resources or restore broad production credentials to make a test succeed.

Record temporary-resource disposition, retained diagnostic access, retention expiry and any continuing charge. Cleanup must identify exact approved objects and owners. If work remains outstanding beyond the declared drain window, leave recovery incomplete and assign the next observation rather than closing the exercise on a green dashboard.

Gate: admission is stopped, baseline behavior is observed, outstanding work reconciles or remains explicitly held, and residual data/resources have an owner. A clean teardown is independent evidence from a passing latency cohort.

12. Hand over an inspectable decision

The service owner receives the frozen contract, fixture hashes, route revision, generator settings, observations, exclusions, class-level results, cost assumptions and stop/recovery receipts. The observer reproduces the counting and percentile method. The owner chooses retain, investigate, authorize a narrower next trial or reject. Meeting a synthetic gate does not authorize production traffic.

The editable offline companion supplies a local worksheet, three synthetic cohorts, arithmetic checks and editable conceptual SVGs. Its checks validate only those supplied records and the displayed arithmetic. They do not contact an endpoint, enforce a firewall or verify a cloud account. Record tested concurrency and arrival schedule when replacing the fixture with authorized observations; the offline example deliberately supplies no throughput or deployment claim.

The probe is done when every class has an accepted result or an owned hold, failures remain in the denominator, attempt/request counts reconcile, cost unknowns are explicit, and restoration has evidence. Next, resolve the largest unknown that could change the decision. If the route passes but data-placement economics remain uncertain, take the packet to the existing placement whitepaper and cost owner rather than expanding this probe into an unsupported migration recommendation.

Related services