Reconcile a Dual Run Before Moving the Write Boundary
Define matching data cuts, classify migration discrepancies and prove recovery before a replacement system receives authority to accept new business writes.
trigger="The replacement is receiving copied data, but the team cannot yet account for missing, changed or delayed business records." owner="The migration owner accountable for the cutover scope and the accepted reconciliation record." participants={['Domain owner', 'Data maintainer', 'Reconciliation operator', 'Security reviewer', 'Cutover operator']} prerequisites={['An approved scope and current write authority', 'Stable business identities and versioned mapping rules', 'A defensible paired data-cut method', 'An isolated fixture and a tested recovery procedure']} outputs={['A versioned invariant register', 'Paired extract manifests', 'An owned discrepancy queue', 'A recovery rehearsal record', 'A scoped cutover decision']} doneWhen={['Every scoped partition has comparable evidence', 'Blocking discrepancies are resolved and rechecked', 'Accepted differences have named owners and expiry', 'In-flight work and external effects are accounted for', 'Write authority and recovery limits are explicit']} />
Establish which records the replacement must preserve
Reconcile a dual run by comparing identified business records at a defensible common boundary, applying versioned rules and accounting for every difference. Keep incomparable data, missing observations and unfinished effects visible. Transfer write authority only after the domain owner accepts the scoped evidence and the operator proves what can be recovered if the replacement fails.
A hypothetical booking platform is moving from a legacy application to a replacement. The old system remains the sole writer while the candidate consumes copied changes. Both contain bookings, cancellations and refund instructions, but the candidate separates records that the legacy schema combines. Equal row counts would not establish that a cancelled booking stayed cancelled or that a refund will happen only once. This procedure tests those relationships without assuming the two schemas should be byte-identical.
The example is an illustrative rehearsal, not an Ampity customer result or a claim that any migration has passed. It assumes synthetic records, inert payment and notification adapters, and one authoritative writer per business partition. Replace those assumptions before execution. Production access, data movement, repair and cutover need explicit authorization; reading this playbook grants none of those permissions.
1. Register the migration scope and current authority
Owner: migration owner with domain owner. Output: scope and authority register. Identify the tenants, business operations, tables, derived views, object stores and downstream effects included in the decision. Name excluded paths rather than letting them disappear from the report. Record whether each partition has a writer, a shadow processor or a read-only candidate. A copied database is not automatically a safe application shadow: a scheduled candidate worker may still send messages or accept updates.
For the booking fixture, include booking identity, assigned resource, cancellation revision and refund obligation. Treat a payment settlement as a separate external record. Identify the system that can establish each fact and the person who accepts its interpretation. Avoid resolving a disagreement by assuming the newer system must be right. A candidate can preserve an old defect, and an approved correction can legitimately produce a different result.
Set the initial exercise boundary to synthetic data and inert recipients. Read back destinations and credential permissions before admitting requests. Record query-rate and storage limits, authorized operators, evidence retention and stop access. If a candidate can reach a live payment processor or the reconciliation account can update arbitrary production rows, hold the rehearsal until those boundaries are corrected and verified.
2. Define business invariants and permitted changes
Owner: domain owner with data maintainer. Output: versioned invariant register. Write assertions about the business relationships, not only individual columns. A booking must belong to the correct tenant; a cancelled booking must not appear as active inventory; an approved refund instruction must retain its amount, currency and source revision. Record the calculation rule, treatment of nulls, precision and historical policy version for each assertion.
Describe deliberate differences separately. The replacement might store a local venue time alongside its zone instead of one formatted string, or replace a legacy display label without changing the resource identity. Specify exactly how the comparator establishes equivalence. Trimming all text, rounding every amount or dropping unknown status values can hide defects. Every normalization should answer a documented representation difference and preserve the underlying business meaning.
Give each rule a reviewer, severity and acceptance condition. In this proposed fixture, cross-tenant records, lost cancellations and duplicate refund obligations block cutover. A display-only difference can be accepted for a stated scope if the owner confirms that no search, integration or contractual output depends on it. These are example decisions, not universal tolerances. Changes to the rules require a new comparison run with its own revision.
3. Establish a comparable data cut
Owner: data maintainer. Output: paired extract manifests. Choose how both systems will represent the same set of committed business changes. In the isolated fixture, stop new admissions to the old writer, settle admitted transactions, identify the last accepted change boundary and wait for the candidate to apply that boundary. Then take the paired extracts while writes remain fenced. Verify that no alternate worker or administrative path can continue writing through the pause.
A continuous comparison needs a different contract: the candidate must expose a demonstrably corresponding state or reproducible input range. A replication timestamp by itself does not prove that two independent queries saw equivalent records. If that contract cannot be established, classify the run as incomparable. Do not take a source snapshot from one moment and target rows from another, then blame every difference on the migration.
PostgreSQL's transaction-isolation documentation explains that Read Committed can use a new snapshot for successive commands, while Repeatable Read retains a transaction snapshot. That can make one extract internally stable; it does not create a shared snapshot across independent databases. Select and test the extraction method for the actual engines and mapping. Long reads also need operational limits so the validation procedure does not damage the system being assessed.
4. Check capture coverage and target application separately
Owner: data maintainer with reconciliation operator. Output: capture and application evidence. Record the initial-load boundary, ongoing change range, retained logs, target application status and any skipped or unsupported objects. Confirm the capture path includes deletes, updates, newly added tables and schema changes relevant to the scope. A healthy task can coexist with an excluded table or an application projection that has not consumed a copied record.
If AWS DMS is involved, its ongoing-replication documentation describes engine-specific change capture and recovery checkpoints. Preserve the selected task's identity and checkpoint evidence. Do not interpret a checkpoint as an application-wide business acceptance certificate. Determine what the chosen task and endpoint configuration establish, including the target's applied state and restart behavior, before using them in the comparison contract.
Test an interruption with synthetic changes while the old writer remains authoritative. After restart, prove that the expected updates and deletes arrived and that no required change range was lost. If retention no longer covers the interruption, stop and plan a controlled reload or other approved recovery. Do not resume from a convenient recent point and report the resulting gap as an accepted difference.
The proposed evidence flow above is not a deployment topology. Arrows carry review artifacts, not permission to repair records or enable writes. The queue retains unresolved and accepted differences. A reviewed decision remains scoped to the recorded boundary, rules and recovery evidence; it does not approve unseen partitions.
5. Pair identities before comparing values
Owner: reconciliation operator with domain owner. Output: identity pairing results. Match records through stable business identities and explicit mappings. Preserve tenant identity as part of the key. A replacement may have new internal identifiers, so an authorized mapping must connect them to the original business object. Reject ambiguous mappings, collisions and orphaned children before comparing display values or totals.
Record expected cardinality for every relationship. One legacy booking may become a booking header and several resource assignments; that is not necessarily a missing-row problem. Conversely, two candidate bookings that point to the same source identity can represent duplication even when the total number of rows looks correct. Check parent ownership and cancellation linkage at the mapped business-object level.
Build partitions small enough to retry and review within the agreed operational budget. Save extract identifiers, mapping revision, object counts, excluded fields and read completeness for every partition. A failed query, incomplete page or inaccessible object becomes unobserved evidence, not a zero count. Protect extracts with restricted access and retention, and avoid copying personal payloads into discrepancy logs when identifiers and redacted values are sufficient.
6. Compare representation and business meaning
Owner: reconciliation operator. Output: reproducible comparison results. Run the approved normalizations and record the resulting rule revision alongside raw evidence references. Compare mapped fields, required relationships and business outputs. For the fixture, verify that a cancellation excludes the booking from active availability and that a refund obligation remains attached to the correct cancelled revision. A matching status column alone cannot establish those derived behaviors.
Keep implementation-dependent values out of equality rules unless the business contract requires them. A generated candidate identifier may differ legitimately; the identity mapping still has to be correct. A money amount cannot be given an arbitrary floating-point tolerance because the schemas use different types. Ask the domain owner to define rounding, scale, currency and the point at which the amount becomes authoritative.
AWS DMS data validation compares source and target rows and distinguishes pending, failed and suspended validation. It also adds query and network load. Use supported row validation as one evidence layer where appropriate, with its limitations recorded. It does not evaluate the booking service's cancellation behavior or certify differently modeled business relationships. Keep tool-level status and domain-level assertions separate in the acceptance record.
7. Classify every discrepancy and keep its denominator
Owner: reconciliation operator with domain reviewer. Output: owned discrepancy queue. Give each difference an identity, affected partition, data-cut references, rule revision, severity, first observation and owner. Distinguish incompatible cuts, missing records, duplicated objects, stale state, value differences and approved representation changes. Preserve the original observation when its classification changes so a later reviewer can reconstruct the decision.
Report populations as well as failures. For example, a synthetic partition may contain twelve mapped bookings, two unpaired records and one lost cancellation. State which assertions ran on which objects and which could not run. Never claim complete coverage from the matched subset alone. A low discrepancy rate across all tenants can still conceal a wholly unobserved tenant or a defect in a rare but consequential operation.
Set ownership and ageing controls for unresolved items. A candidate record that is merely waiting for a documented apply boundary can be rechecked after that boundary; it cannot remain indefinitely excused as replication lag. Accepted differences need the approving owner, reason, affected scope and review expiry. An AI assistant may group redacted discrepancy descriptions for human triage, but it must not invent missing evidence, approve a tolerance or close a blocking item.
8. Repair only through an approved, version-checked path
Owner: data maintainer with domain reviewer. Output: repair receipts and revalidation. Establish whether a defect belongs in extraction, mapping, capture, candidate processing or the authoritative data. Correct the responsible layer through a reviewed change. Avoid patching a target row solely to make the comparison green while the same defect remains in the transformation or replay path. Preserve the failing fixture so the corrected revision can be tested again.
For any approved data repair, identify the exact record revision expected before writing and verify the authoritative source immediately beforehand. If the record has changed, hold the repair and re-evaluate it. A stale overwrite can undo a legitimate cancellation or revive a deleted object. Use the actual system's concurrency and audit mechanism; this procedure does not prescribe a generic SQL update against production.
Read back the repair and rerun the affected assertion plus its dependent relationships. Check the operation's external effects separately. If a repair request timed out after submission, treat its result as unknown until authoritative evidence resolves it. Blindly retrying may duplicate an obligation. Keep failed, rejected and uncertain attempts in the queue rather than marking the item repaired from an acknowledgement alone.
9. Settle external effects and unfinished work
Owner: domain owner with service maintainer. Output: operation settlement ledger. List the admitted operations that can outlive the data copy: queued notifications, refund instructions, reservations and scheduled updates. Record their operation identity, owning system, input revision and confirmed external disposition. In the fixture, inert receivers allow a replay to be inspected without sending a real notification or payment. Prove those receivers are still inert when delayed work resumes.
AWS's transactional-outbox guidance discusses consistency between a database change and its emitted event, including duplicate delivery and idempotent consumption. An outbox can support a design; it does not establish that every migrated consumer or external recipient deduplicates correctly. Test the same operation through interruption and replay, retaining recipient evidence rather than counting successful worker attempts.
Assign each unfinished operation to one authority before transfer. Do not let old and new workers both claim a refund after cutover. Stop conditions include unexplained duplicate effects, lost ownership and unknown outcomes that the chosen recovery path cannot settle. The migration owner must know whether the residual work will be completed, held for investigation or explicitly cancelled under an approved business procedure.
10. Rehearse the cutover fence and recovery limit
Owner: cutover operator with migration owner. Output: authority-transfer rehearsal. In isolation, fence old admissions and every alternative write path, settle in-flight work, record the terminal accepted boundary and verify candidate application through that boundary. Run the final reconciliation against the stable scope. Transfer authority to the candidate only after the reviewer accepts the record. Test that stale old workers, scheduled jobs and retried requests cannot continue writing.
Distinguish recovery before and after the first candidate write. Before that write, returning authority may be possible if the old system is still intact and no external effects escaped. After candidate writes begin, a routing change alone can discard new records or replay effects. Rehearse how new operations are identified, how compatible state reaches the old system or an approved alternative, and which writes require a hold instead of automatic reversal.
Record the point beyond which the proposed reversal is unsupported. Do not promise bidirectional safety from two replication tasks or assume a new schema can be translated backward without loss. If reverse reconciliation is unavailable, the plan may require a bounded write pause and forward repair. Exercise that decision with a synthetic candidate-only update and a delayed external receipt before approving the runbook.
11. Assemble the acceptance record and hold criteria
Owner: migration owner with domain reviewer. Output: signed scope decision. Link the register, paired manifests, comparator revision, complete discrepancy queue, settlement ledger and recovery rehearsal. State the actual covered tenants, operations and boundary. Include the known limitations and the time at which evidence was obtained. A dashboard that says passed cannot replace that record when some partitions were skipped or the rules changed after the run.
The following acceptance criteria are a proposed review checklist for the synthetic booking rehearsal. Adapt them to the real business contract and data risks. They do not authorize a production cutover or certify a database product.
- Every scoped partition has complete extracts at the accepted comparison boundary, or remains an explicit hold.
- Identity mappings are unambiguous, tenant-correct and consistent with the approved cardinalities.
- Blocking invariants pass; every accepted difference has an owner, scope, reason and review expiry.
- Missing, pending, suspended and incomparable evidence is counted separately from a passing result.
- External effects and unfinished operations have one accountable authority and an evidenced disposition.
- The write fence and recovery limit are rehearsed, including a candidate-only write and an unknown external result.
Hold the transfer if any prerequisite loses validity, if a blocker is reopened, if a repair lacks readback or if evidence becomes incomplete. The reviewer can approve a smaller demonstrably isolated scope, but must record how that scope avoids unresolved dependencies. No deadline or aggregate pass percentage substitutes for accepting a lost cancellation or cross-tenant record.
12. Verify the new authority before retiring the old path
Owner: migration owner with service maintainer. Output: post-transfer review and retirement conditions. After an authorized transfer, check fresh operations through the candidate writer and their authoritative readback. Continue the agreed reconciliation for the observation period, including late arrivals, deletes and retries. Record old-writer rejection and downstream ownership. Choose the observation period from the actual operation lifecycle, not an arbitrary number of hours that excludes the weekly settlement job.
Keep the old evidence, mappings and recovery access for the approved retention period without leaving unintended writers active. Retiring the old path is a separate decision: resolve residual work, verify restored data if needed, revoke obsolete access and account for retained personal data. A stopped replication task is not evidence that every old credential, scheduled consumer or diagnostic extract has been removed.
For related decision framing, read Compare Old and New Systems Before Changing Authority. For reversal hazards beyond routing, use the feature-flag rollback discussion. Ampity's platform modernization and system architecture design services can help define the comparison and recovery boundary. Reading or downloading does not require an email; contact remains your choice.
Start the next migration review with one paired fixture, its invariant register and every unresolved discrepancy. The migration owner should reject an unpaired or incompletely observed result before discussing a wider cutover.