Rehearse DNS Failover and Client Recovery

Measure DNS failover through resolver caches, connection reuse and an accepted business operation. Keep authoritative propagation, client recovery and safe reversal...

trigger="The recovery plan switches a DNS destination, but the team has not observed how representative clients leave the old endpoint and complete their business operation." owner="The service owner accountable for accepting client-visible recovery." participants={['DNS operator', 'Application engineer', 'Network observer', 'Data owner', 'Independent exercise observer']} prerequisites={['Approved isolated hostname and bounded exercise window', 'Verified primary and recovery targets with distinguishable identities', 'Synthetic business fixtures and blocked external effects', 'Documented client cohorts, stop control and reversal authority']} outputs={['A client-cohort and dependency register', 'Authoritative DNS and resolver observations', 'Connection and business-operation timing results', 'An uncertain-write reconciliation record', 'A tested reversal and an owned defect list']} doneWhen={['Required client cohorts pass their agreed business check', 'DNS and connection milestones remain separately measured', 'Unsettled effects are reconciled or explicitly held', 'Reversal is verified through representative clients', 'The service owner accepts results and stated coverage limits']} />

Measure the client-visible result

Use this playbook when a recovery procedure depends on DNS routing and the team needs evidence that its clients actually resume service. Record the authoritative routing change, observations from selected resolvers, connection behaviour and completion of an agreed business operation separately. A successful DNS change is an intermediate milestone. It cannot tell you whether a long-running client has opened a new connection or whether the recovery application has the state needed to finish its work.

Start with an isolated hostname and a bounded service function. A synthetic appointment lookup and an approved test update provide clearer acceptance evidence than a generic homepage request. Give the primary and recovery targets distinguishable identities that can be observed without exposing internal secrets. Establish how the observer knows which target processed a request, including any proxy or load-balancing layers that affect that observation.

This is proposed engineering guidance. The appointment example, client cohorts and evidence model are illustrative, not an Ampity customer result or a DNS-provider recovery guarantee. It does not authorize production routing changes, deliberate public DNS disruption or replaying customer writes. A platform operator must supply the provider-appropriate execution instructions and obtain the approvals needed for the actual environment.

The DNS caching article explains the mechanisms behind delayed client recovery. The unsettled-write article covers effects that can remain ambiguous during a switch. This rehearsal produces observations and assigned follow-up work for one stated scope.

The three cohorts are independent observations, not a required sequence or a claim that all clients behave alike. The DNS observation supplies routing evidence; it does not force a cached resolver or an existing connection to change destination. The owner closes the exercise using each required cohort's results.

1. Agree scope, acceptance and stop conditions

Owner: service owner. Output: approved exercise contract. Name the hostname, record types, routing mechanism, participating clients, permitted operations and observation window. Record the change approver and the person who can stop or reverse the exercise. Keep the test hostname separate from public service names unless a separately approved production exercise is explicitly intended.

Define the business check before selecting a failure trigger. For the illustrative appointment service, a required client might need to fetch a known appointment, confirm its current revision and perform an isolated status transition under its normal identity. The result must preserve the tenant boundary and expected revision. A health page returning 200 is insufficient for those requirements, though it can help localize a later failure.

Specify stop conditions, including an unexpected production destination, uncontrolled writes, exposed sensitive payloads, ambiguous target identity, loss of the stop control or service deterioration outside the approved scope. Assign the response to each condition. A failed drill can produce useful evidence, but it remains held until the unsafe or unexplained condition has an owned disposition.

Agree the time boundary. Measure from the approved trigger or simulated failure declaration through acceptance of the required cohorts. Keep detection, approval and execution delays visible when they are part of the recovery objective. Do not start the clock only when the DNS change has already finished, then present the shorter interval as end-to-end recovery.

2. Map the resolution and traffic path

Owner: DNS operator with network observer. Output: dependency register. Identify the authoritative zone, record chain, applicable routing policy, health signal and configured TTLs. Include aliases or CNAME dependencies where relevant, and IPv4 and IPv6 paths the service actually uses. A test observing only one record family cannot establish recovery for another family still pointing elsewhere.

RFC 1035 describes the TTL as a cache-time interval for a resource record. Use the record TTL and observed remaining TTL as inputs to the test, not as a promise that all applications will recover in that interval. Keep the answer, resolver identity and query time with each observation so later comparisons do not mix different cache states.

Map the network path as well. A client may use a local resolver, enterprise resolver, VPN, HTTP proxy or another connection intermediary. The application can also keep its own address or connection state. State which layer your observation inspects. A lookup from an operator's laptop using a public resolver is evidence about that path, not every customer's runtime.

If TTL values are lowered for a planned exercise, allow for answers cached under the earlier values. Record when the preparation change was made and which cache states the test intentionally preserves. Lowering the authoritative value immediately before the failover does not demonstrate that previously cached answers have been invalidated.

3. Prove the recovery target before changing routing

Owner: application engineer with data owner. Output: target-readiness evidence. Verify the recovery target's identity, certificate, hostname handling, authentication, application build and required dependencies. Test it through an approved direct-target diagnostic path before changing the test hostname. Keep the hostname's TLS and application routing requirements intact; disabling certificate verification can conceal a target-readiness defect.

The curl manual documents --resolve as supplying an address for a host and port instead of normal resolution. Such a test can help isolate endpoint readiness, but it intentionally bypasses the ordinary DNS answer path. Label its results as direct-target evidence. Do not include it in the resolver-recovery measurement or use it to conclude that clients have observed the DNS change.

Validate the selected synthetic business function at the recovery target. Confirm the required records, schema and permissions, not merely a listening port. Block real payments, notifications and other external effects before enabling test writes. If data readiness is outside the rehearsal scope, use read-only acceptance and state that write recovery has not been established.

Check capacity and dependency assumptions appropriate to the bounded test. A lightly loaded isolated target can prove a connection path without establishing production capacity. Record that limitation rather than injecting unapproved load. Hold the routing exercise if the target cannot perform its required function through the controlled diagnostic path.

4. Prepare representative client cohorts

Owner: application engineer. Output: cohort register and test harness. Prepare a fresh-lookup cohort, a warm-resolver cohort and an existing-connection cohort where those states occur in the real workload. Name each runtime, version, network path, resolver, connection policy and intended test operation. A mobile application, browser and server-side worker may require different harnesses rather than three copies of the same command.

Warm the selected cache or connection deliberately and verify the state before the trigger. Record the answer and remaining TTL seen by the warm-resolver cohort. For an existing connection, record its established destination using an approved observation mechanism. Calling a process warm because it has been running for an hour does not prove that it retains the connection under test.

Use client-specific instrumentation for reuse. The curl manual explains that connection reuse occurs across transfers in a single invocation, not across separate curl runs. A loop launching a new process for every request therefore cannot represent a persistent client's connection lifecycle merely because the loop itself is long-running. Record what the harness preserves and what it resets.

Keep authentication and synthetic state stable across cohorts unless their restoration is intentionally part of the drill. Test permitted and denied identities where the business function requires it. Limit the conclusion to the covered clients and network paths. If a relevant runtime cannot be observed safely, mark it untested and assign the missing harness as follow-up work.

5. Capture the baseline and observation clock

Owner: independent observer. Output: baseline evidence packet. Before the trigger, collect authoritative answers, selected recursive answers, target identities and business-check results from each cohort. Retain explicit timezones and synchronized observation sources. Record any expected delay or uncertainty in timestamps so a small apparent ordering difference is not presented as a measured system guarantee.

Use a stable synthetic operation identifier to correlate the application result with target-side evidence. Keep credentials, personal data and request bodies out of broad logs and exported reports. If an internal target label is sensitive, use an approved restricted record and a non-sensitive label in the shared report. The label must still allow the observer to distinguish the destinations.

Capture baseline failures rather than suppressing them. If a cohort already cannot complete the accepted operation, the exercise cannot attribute its later failure to the routing event without more investigation. Correct the baseline or record the result as inconclusive. Keep DNS, connection, authentication and business errors as separate classes in the harness.

State the sampling interval and evidence retention. A probe every thirty seconds cannot establish the precise instant of recovery within that interval. Report the interval containing the transition, or an upper bound based on the observed samples. Avoid inventing millisecond precision from a coarse polling log.

6. Execute one approved routing or failure trigger

Owner: DNS operator with exercise commander. Output: trigger record. Apply only the reviewed trigger to the confirmed test target. It might be a controlled routing change or an approved health-signal change. Record which mechanism was used. A manual record update and a health-check-driven routing decision are different rehearsals; one does not automatically prove the other works.

Capture the change identifier, submission time and provider-side result. For Route 53 change batches, AWS's GetChange reference defines INSYNC as propagation to the Route 53 DNS servers managing the hosted zone. Treat that result as authoritative propagation evidence, not an assertion about external resolver caches, client connections or business completion.

Observe the agreed authoritative answers independently where practical, with the record chain and target labels retained. Then keep watching the client cohorts without resetting their warmed state. Clearing every cache at the trigger would answer a different question: how a freshly initialized client behaves, not how existing clients recover.

Do not make successive unplanned changes while waiting for a cohort. Repeated flips can produce a mixed set of cached answers and obscure which trigger caused a transition. If a stop condition is reached, hold the test and use the approved recovery procedure. Record a revised attempt separately rather than merging its observations into the original timeline.

7. Observe resolvers without assuming universal expiry

Owner: network observer. Output: resolver observation log. Record the query source, resolver, answer, remaining TTL, timestamp and result category. Keep authoritative observations separate from recursive results. When a recursive answer changes, retain enough evidence to show the target distinction rather than logging only a successful query exit code.

RFC 8767 describes serve-stale behaviour for recursive resolvers when authoritative data cannot be refreshed. It updates the TTL interpretation for that exceptional case. This does not mean every resolver serves stale answers or that every failover encounters it. Record whether the tested resolver supports it, whether its refresh path is available and what behaviour was actually observed.

If the exercise includes resolver refresh failure, simulate it only inside an approved isolated resolver setup. Do not disrupt public nameservers or a shared enterprise resolver to make a demonstration more realistic. Use the result to describe that configured path and its limitations. A fresh public-resolver probe cannot establish the behaviour of the isolated serve-stale scenario.

Preserve mixed outcomes. One cohort may resolve the recovery target while another still sees the primary. That evidence identifies a recovery gap, even if the configured TTL is short. Do not force all results into one propagation-time number or declare the slower cohort irrelevant after the exercise has started.

8. Check connections and business effects together

Owner: application engineer with data owner. Output: client recovery and effect record. For each cohort, observe the endpoint that processes its requests, connection renewal or failure, retries and completion of the accepted business check. Keep the first new DNS answer, first connection to the recovery target and first accepted operation as separate milestones. Their ordering and elapsed intervals explain different constraints.

An existing connection can remain useful or can fail after the DNS route has changed. Document the intended treatment: continue safely, drain, reconnect or hold. Test the actual client policy under the approved conditions. Do not assume that a DNS lookup forcibly moves an established connection, or that every successful old-target request constitutes a failed recovery.

For test writes, reconcile uncertain outcomes before retrying. In the appointment example, a timed-out status update may have committed at the primary. A retry at the recovery target must not create a duplicate transition or overwrite a newer revision. Use the application's supported operation identity and authoritative state check. If those controls are absent, hold writes and record the recovery limitation.

Keep external effects visible. A success response does not establish that a notification or downstream update occurred exactly once. The isolated rehearsal should use blocked or synthetic destinations, with a ledger of intended and observed effects. Do not replay live requests to manufacture confidence about effects the exercise cannot safely inspect.

9. Rehearse reversal through the same client states

Owner: exercise commander with DNS operator. Output: reversal evidence. Before reversing, confirm that the original target can safely resume the selected function and that data ownership is unambiguous. A DNS change back cannot repair conflicting writes. The data owner must accept reconciliation or keep writes held before routing is restored.

Execute the approved reversal and retain its own change identifier and timestamps. Repeat authoritative, resolver, connection and business checks for the required cohorts. Leave cache and connection state intact where the test intends to observe it. The reverse change can encounter the same delayed or mixed client states as the forward change.

Record a failure path when reversal is not safe. The agreed fallback may hold the test endpoint, keep the recovery target serving a restricted function or stop synthetic traffic while a defect is investigated. Define that decision before the drill. Automatically flipping back when an alarm fires can reintroduce an unhealthy target or hide unresolved effects.

Do not delete the forward attempt after a clean reversal. Keep failures, manual interventions and elapsed intervals in the report. Remove temporary records, identities and test resources through approved cleanup steps only after evidence is retained and remaining owners have accepted their tasks.

10. Close the acceptance record and schedule the next drill

Owner: service owner with independent observer. Output: accepted, held or inconclusive result. Use these acceptance criteria for the stated scope. Retain evidence references and owners next to each result rather than making one undifferentiated pass statement.

  • [ ] The trigger, routing mechanism and affected hostname match the approved contract.
  • [ ] Authoritative propagation and recursive observations remain separately recorded.
  • [ ] Every required cohort has an observed target and a business-check result.
  • [ ] Warmed state was preserved or every deliberate reset is identified.
  • [ ] Uncertain writes and downstream effects have an owned disposition.
  • [ ] Reversal or the approved hold path was observed and accepted.
  • [ ] Timing includes the declared boundary, sample interval and failed attempts.
  • [ ] Untested clients, capacity assumptions and evidence gaps remain visible.

For each cohort, record the trigger time, authoritative milestone, first observed recovery answer, first observed recovery connection and business acceptance time. Mark a milestone not observed when the harness cannot establish it. Calculate elapsed time from the same declared start, and do not add overlapping intervals as if they were sequential work. Keep the raw restricted log available for review.

Assign defects by mechanism: resolver freshness, client connection lifecycle, target readiness, authentication, data state or uncertain effects. A TTL change cannot fix every category. Link each proposed correction to a repeated test that would prove the intended improvement, with its owner and required approval.

The next action is to rehearse one hostname and one bounded business function with the three relevant client states. Bring the cohort register and unresolved evidence to a reliability review. Repeat the drill after a consequential change to DNS policy, client runtime, connection settings or recovery dependencies, and keep the conclusion limited to what the exercise observed.