DNS Failover Is Not Instant Client Recovery
Test endpoint changes across authoritative DNS, recursive caches, runtime resolution and connection reuse. Measure client recovery instead of assuming TTL is an SLA.
A changed answer is only one part of recovery
DNS failover changes where a client can resolve a service name. It does not prove that every client has obtained the new answer, opened a replacement connection or completed a useful operation. Test authoritative answers, recursive-cache behavior, application resolution and connection reuse as separate boundaries. A record's TTL is not an end-to-end recovery promise.
The incident often appears inconsistent. An engineer's fresh command reaches the replacement service, while a long-running worker still fails. Some customers recover before others. The DNS console looks correct, but the team cannot say which clients have switched or whether the replacement handles the same authenticated workflow.
This article proposes an endpoint-change rehearsal for service owners. Examples and the evidence record are illustrative, not measured customer outcomes. It covers client recovery after a destination change, not the complete design of a global traffic manager or a guarantee that a particular resolver behaves identically everywhere.
Identify which layer supplied the answer
An authoritative DNS answer describes the record at its source. A recursive resolver may answer from its cache. The application may use another resolution path or retain an address locally. A transport connection can then outlive the lookup that established it. These observations concern different layers and should not be collapsed into one “DNS works” result.
For ordinary simple records, Route 53's TTL guidance describes the cache interval requested of recursive resolvers and the tradeoff between longer caching and slower adoption of changed records. That setting does not specify the lifetime of an already-established application connection.
Inventory the actual path for each important client cohort. Identify runtime, version, operating environment, configured resolver and any proxy that performs resolution on the application's behalf. Record the effective service name, including relevant aliases, without assuming every client resolves the same name through the same chain.
Use the inventory to place probes. A command run on an engineer's laptop is useful for that laptop's path. It is not evidence for a container's resolver or a managed worker's runtime. Label each observation with the vantage point so apparently conflicting results become understandable instead of being discarded as noise.
A lower TTL does not retroactively shorten an earlier cached answer
Consider an illustrative simple-record change. A resolver caches an answer carrying a 300-second TTL immediately before the team lowers the authoritative record's TTL to 60 seconds. The team then changes the destination and expects every client to switch within one minute. The earlier cached answer did not acquire the new TTL merely because the console setting changed.
This example explains a dependency, not a universal convergence calculation. Resolver policy, alias chains and the application's resolution behavior can change the observations. Record when the earlier answer was obtained and its remaining TTL if that information is available. Do not manufacture a cache-age estimate from the time the operator saved the record.
For a planned change, reduce the relevant TTL sufficiently before cutover to allow previous cached answers to age out, subject to the chosen DNS mechanism and verified client behavior. Confirm adoption of the shorter setting from representative paths before treating it as an operating assumption.
For an unplanned failure, that preparation window may not exist. The recovery design must tolerate clients following the previous answer or failing to establish a useful connection while caches converge. A lower configured TTL may be helpful, but it cannot remove a dependency that is already present in the client fleet.
Include exceptional resolver behavior without assuming it is universal
RFC 8767 defines serve-stale behavior that can allow a recursive resolver to use expired data when it cannot refresh from authoritative servers. This is a resilience mechanism under a refresh failure, not evidence that every resolver ignores TTL or that ordinary successful refreshes always return stale answers.
That distinction matters during compound incidents. A destination failure and an authoritative-resolution problem may coexist. If a resolver keeps supplying a previous destination under its configured stale-answer policy, the client may behave differently from the team's ordinary endpoint-change rehearsal.
Where you operate the resolver, record the actual configuration and test the approved refresh-failure scenario separately. Where it belongs to another operator, mark the behavior as an external dependency and state which observations you can obtain. Do not claim to control public or customer-managed resolver policies.
Preserve the response status and observation time alongside the answer. An older destination returned from a resolver is not, by itself, proof that serve-stale caused it. It could still be an unexpired cache entry, another alias, or a different name-resolution path. Diagnose from evidence rather than naming the most sophisticated mechanism first.
Connection reuse can make a fresh lookup irrelevant to existing traffic
An application does not necessarily resolve a hostname for every business request. A retained transport session or connection pool can continue using the destination chosen earlier. A successful fresh lookup therefore cannot establish that existing traffic has moved to the replacement.
Rehearse both new and long-lived clients. Observe the destination of an actual connection through approved telemetry where available, rather than inferring it from a lookup performed by a separate tool. If a proxy owns the connection, observe the proxy's path as well as the client-facing connection.
Choose a connection-refresh policy deliberately. Shortening every connection lifetime can increase setup work; restarting every client can discard useful state or create a reconnection surge. Neither is a universal fix. The policy must fit the protocol, deployment model and operation deadlines, and it needs a tested response when the old destination stops accepting work.
Do not bypass hostname verification by hard-coding an IP address as a casual workaround. Recovery must preserve the service's TLS identity, authentication and authorization expectations. A connection to a reachable machine is not equivalent to a safe connection to the intended service.
For database clients, the companion RDS connection-pool failover review covers unusable sessions and interrupted transactions. The DNS review establishes where connections go; it does not resolve whether a business write completed before connectivity changed.
Build a rehearsal around observable client cohorts
Use an isolated environment with two distinguishable destinations and synthetic traffic. Ensure both destinations have the intended service identity and safe test data. Assign a test-window owner, a rollback boundary and an explicit prohibition on uncontrolled customer traffic or external charges.
Define cohorts that expose the likely disagreement: a fresh process, a warm long-running process, an existing pooled connection and a newly created connection from that same application. Add important background workers or customer runtime types where they are genuinely part of the service's recovery commitment.
Confirm a baseline for each cohort before changing the endpoint. Then record authoritative answers, selected recursive answers, application connection attempts and the result of the required operation. Use a common clock reference and identify any timestamp uncertainty. The experiment should reveal the ordering of observations without pretending all components have perfectly synchronized clocks.
After the change, continue observing long enough to see the chosen acceptance condition or an explicit test-budget failure. Do not stop at the first successful fresh client. Record cohorts that require a manual intervention and preserve the intervention's time so it is not counted as spontaneous recovery.
The required operation should include the application boundary that matters. For example, a synthetic authenticated read can verify more than a TCP handshake. A write rehearsal additionally needs operation identity and reconciliation. State exactly which capability was tested rather than calling every probe a full recovery check.
Use a record that keeps the layers separate
The proposed record below deliberately leaves results pending. Its purpose is to make missing evidence visible, not to supply a completed failover report or production configuration.
scenario: "Endpoint change test"
name: "Synthetic service name"
cohort: "Warm background worker"
before:
answer: "Previous destination"
ttl_seen: "Pending observation"
cache_age: "Unknown"
after:
authoritative: "Pending"
recursive: "Pending"
runtime_path: "Pending"
fresh_connection: "Pending"
reused_connection: "Pending"
operation:
identity_check: "Pending"
authenticated_read: "Pending"
intervention: "None recorded"
decision: "Not accepted yet"Keep one record per meaningful cohort and correlate it with the test window. Store restricted telemetry references rather than publishing internal hostnames, customer identifiers or security-sensitive network details. Mark unavailable observations as unavailable; do not substitute a different layer's success just to complete the form.
Summarize recovery using the cohort results and the team's predefined objective. Distinguish a first success from sustained success, and a tested population from all possible clients. If one important cohort remains on an unusable path, that is an unresolved recovery dependency even when the DNS management interface reports the desired record.
Fix the layer that failed, not the nearest visible setting
Start with one warm application process and one fresh process in an isolated environment. Record their actual resolution paths, change the synthetic service destination, and compare connection destinations and authenticated-operation results. Use that evidence to select the next cohort instead of assuming the first successful lookup represents everyone.
If authoritative answers are wrong, investigate the record change and routing mechanism. If recursive answers differ, inspect the observed cache and resolution path. If fresh connections recover but reused ones do not, inspect connection ownership and refresh behavior. If both reach the new destination but the workflow fails, examine the replacement service's readiness and application contract.
These are investigation branches, not automatic diagnoses. Preserve the evidence that selected the branch and repeat the same cohort test after the correction. A smaller TTL cannot fix missing permissions at the replacement, and a process restart cannot prove that the authoritative record was correct.
Document what remains outside your control, including customer runtimes or third-party resolution behavior. Use those dependencies to define honest recovery expectations and supported procedures. Ampity's reliability review can help scope an application-specific rehearsal, but neither this proposed method nor a green lookup warrants an instant-failover claim.