Data Reconciliation Before Modernization Cutover
Choose defensible data cuts, business invariants and divergence recovery before a replacement platform gains write authority. Includes comparison fixtures.
audience="Teams replacing a live application while bookings, payments, cancellations or other business records continue to change." decision="Choose what agreement means, how both systems become comparable and which evidence permits a scoped transfer of write authority." position="Compare independently defined business invariants at defensible paired data cuts. Keep freshness, correctness, coverage and unsettled effects separate. A passing copy check alone does not authorize cutover." scope="A proposed decision framework with synthetic booking records and inert downstream adapters, not a customer migration result or an instruction to repair production data." outputs={['An invariant register', 'A comparison-method decision', 'Paired extract manifests', 'An owned divergence record', 'Independent fixtures', 'A fenced cutover and recovery contract']} />
Executive summary
Data reconciliation establishes whether the replacement preserves the business meaning that a modernization must protect. It compares identified records at a defensible common boundary, applies declared mappings and explains every material difference. It does not mean taking two live exports at approximately the same time and celebrating equal totals. Nor does it mean expecting different schemas to contain identical rows. A useful acceptance statement names the population, versions, observations, unresolved effects and decision owner.
The central architectural choice is how to make observations comparable while the business keeps changing. A controlled write pause simplifies that argument but may disrupt customers. A consistent snapshot followed by a verified change stream can reduce the pause, at the cost of continuity and transformation controls. Online shadow comparisons reveal behavior under real requests but cannot, on their own, establish agreement for unvisited historical records. Many migrations need a combination. Each method has a different blind spot, so do not combine their percentages into one unqualified success score.
For this paper, a hypothetical booking platform moves from a combined legacy record to separate reservation, cancellation and refund-intent records. The old application initially remains the only authorized writer. The candidate consumes copied changes and performs no real payment or notification actions. This is a synthetic engineering scenario. The example amounts, identifiers and fixture expectations illustrate a review method; they are not Ampity delivery results, customer data or migration performance measurements.
The recommendation is to separate four decisions: what must remain true, what evidence makes the observations comparable, what a discrepancy requires and when authority may move. Preserve unknown states instead of rounding them into agreement. A replacement that matches the old database can still reproduce an old defect. A replacement that looks correct can still lack required historical data. And a replacement that was correct before cutover may require a new recovery strategy after it accepts its first business write.
Define agreement as a business contract
Begin with the decisions that customers and operators rely on. A confirmed reservation occupies a particular slot for a particular tenant. A cancelled reservation must not continue reserving capacity under the declared cancellation policy. A refund intent must retain its relation to the original charge and must not create another refund merely because a copied event is replayed. These are business invariants. Row count, checksum and replication lag are supporting observations, not substitutes for the invariant.
Write each invariant with a stable business identity, included state, comparison rule, allowed representation differences, blocking conditions and accountable owner. The booking identifier should be qualified by tenant if identifiers are not globally unique. The comparison contract should define whether a deleted record must be absent, represented by an authorized retention record or excluded under an approved scope. Do not infer that any absent row is harmless just because the replacement uses a different lifecycle model.
Distinguish preservation from correction. If the legacy system contains a reservation that violates the approved capacity rule, faithful copying and a valid future business state are different goals. Record the defect, the authority to change it and the migration treatment. A reviewer may choose to preserve historical evidence while correcting the operational projection. That decision should be explicit and testable. Silently editing the source or normalization logic until both systems agree destroys the evidence needed to understand what happened.
Specify independently expected outcomes before running candidate code. The expected cancelled state should come from the accepted policy and known synthetic input history, not from whichever system currently looks more plausible. The old application is the current writer in this scenario, but that fact does not make every old output a reliable oracle. Agreement between two implementations is useful evidence; independent invariant evaluation is what exposes a shared misunderstanding.
Choose a comparison method for the actual constraints
A paused comparison is appropriate when the organization can stop every scoped writer, settle in-flight transactions and capture stable observations within its accepted disruption window. Its attraction is a simpler evidence argument. Its cost includes customer delay, scheduled jobs, partner callbacks and operator preparation. A maintenance banner does not prove that writes stopped. If a back-office import or delayed payment event can still change the population, the advertised pause does not create the boundary required by the comparison.
A snapshot-and-change-stream method is useful when a long pause is unacceptable. It requires an explained relationship between the initial state and subsequent committed changes. The team must establish where capture starts, whether retained history covers the required interval and what position the candidate has actually applied. A consumer reporting that it received an event is weaker evidence than a durable application record demonstrating its committed effect. Transformation and replay behavior remain application responsibilities.
An online shadow comparison answers a narrower question: do old and candidate interpretations of selected requests agree under comparable input and policy? It can expose hot-path differences early without handing the candidate business authority. It can miss dormant tenants, historical cancellations and unpopular query paths. Sampling should be intentional, with coverage gaps visible. A shadow must not send a second email, charge or refund. Copying live requests is not permission to duplicate downstream actions.
A combined design may use row validation for copied storage, invariant checks for remodeled entities and shadow reads for user-visible behavior. Report each method separately, with its included population and exclusions. Choose the simplest combination that proves the accepted contract, rather than adding mechanisms because the topology appears more sophisticated. If no defensible paired cut is available, record the comparison as inconclusive. More queries cannot repair an observation boundary that was never established.
Establish paired data cuts, not matching wall clocks
A data cut is a declared set of committed business states included in an observation. It needs enough evidence for a reviewer to determine which changes belong inside it. Extract timestamps are useful operational metadata, but they are not automatically commit boundaries. Two machines can show the same time while representing different transactions. An export can also span changes made during its execution. Describe the database isolation, extraction method, included partitions and completion evidence instead of calling a timestamp a snapshot.
PostgreSQL documents that Read Committed can expose different snapshots to successive commands, whereas Repeatable Read retains the transaction's snapshot across successive reads. This supports a local extraction choice, not a shared snapshot across independent databases. The application must still explain how the source cut corresponds to the target's applied state. See the official transaction isolation documentation. Recheck behavior for the actual engine and version before selecting an extraction procedure.
In the proposed comparison, each source and candidate manifest includes the scope revision, extraction identity, key ranges, schema and mapping versions, cut evidence, record counts, failures and artifacts. The pairing record explains why both observations contain the same committed population. If the target is ahead, ordinary current-state reads may no longer represent the earlier source cut. Use a defensible retained observation, a supported as-of method or a new pair. Do not describe a merely advanced checkpoint as equality at an earlier point.
Keep freshness separate from comparability. A candidate may be fresh but incorrectly mapped. A correct candidate observation may be too old for the proposed cutover decision. Declare an acceptable age and the invalidating changes that require another review. Data, mapping, policy and writer configuration can all change after a passing run. The evidence should remain bound to the reviewed versions, not travel indefinitely as a green badge attached to the project.
Follow lineage through the invariant map
The first visual shows why a replica comparison and a business comparison answer different questions. The legacy observation and candidate observation each feed a declared projection. The reviewer compares those projections under an independently defined invariant register and preserves the result with both manifests. The neutral components represent responsibilities, not a selected cloud product or an implemented deployment. No connector in this map grants repair or cutover authority.
An invariant can depend on more than one storage record. A reservation's effective capacity state may combine its latest revision, cancellation policy and tenant calendar. The comparison should disclose those dependencies. If the projection reads a current calendar while evaluating an older reservation cut, it may apply a rule that was not in force for that decision. Bind required policy and reference data to the observation or explain the accepted temporal interpretation.
Lineage also makes omissions visible. A replacement might preserve reservations while failing to copy their attachments, audit evidence or unresolved refund intent. Those assets should not disappear merely because they sit outside the main relational database. Identify each included store and consumer. Where the review intentionally excludes an analytical view or historic binary, name the owner accepting that exclusion and the consequence. A narrow claim with explicit omissions is more credible than a complete-looking map with hidden dependencies.
Normalize representation without normalizing away defects
Use a versioned mapping to turn different schemas into comparable business facts. A combined legacy booking may become a reservation plus a cancellation record. Different physical layouts are not a discrepancy when they preserve the accepted lifecycle. The projection should retain the original business key, relevant revision, state, amounts and links to supporting evidence. Keep the raw observation available under the approved retention policy so a reviewer can distinguish a source defect from a projection defect.
Define null, missing and empty values separately wherever their meaning differs. A missing cancellation time should not silently become the current time. Currency comparisons should preserve currency identity and the declared minor-unit convention. Time comparisons need an explicit timezone and daylight-saving interpretation. Rounding should be stated at the business boundary that actually permits it. Using a wide numeric tolerance to absorb unexplained money differences turns a detection mechanism into a defect-hiding mechanism.
Allowed differences need owners and reasons. A display label may change intentionally without altering settlement obligations. A historical timestamp may be represented with a different precision while preserving the accepted ordering rule. Record those cases before evaluation, with a narrow scope and expiry if temporary. A tolerance for one derived display field must not leak into the reservation identity, authorization state or refund amount. Do not enlarge tolerances after discovering a failing fixture merely to protect a release date.
Test the projection as software. Feed missing keys, duplicate identities, unexpected enum values, tenant collisions and schema changes into the mapping. Require explicit failure rather than silently dropping records that cannot be parsed. A mapper that excludes difficult rows can create excellent-looking agreement among the survivors. Include rejected and unobserved records in the intended population, and retain their disposition. Successful comparison of a reduced subset is not successful comparison of the declared scope.
Account for change continuity, deletes and transaction boundaries
Change capture reduces the distance between the initial copy and current state, but receiving changes does not prove the application's reconstructed business history is complete. Record the stream's origin, retention assumptions, consumer state and applied position. Demonstrate how the initial snapshot joins the stream without an unexplained gap or duplicate effect. If the required history has expired, restart from an authorized baseline rather than claiming continuity from a checkpoint that no longer has supporting source history.
Debezium's PostgreSQL connector documents its initial snapshot and subsequent committed insert, update and delete events. Its snapshot behavior is configurable, and consumers still need an explicit handling contract. The official PostgreSQL connector documentation also describes delete and tombstone events. Check the selected connector and transformations rather than assuming downstream records automatically disappear. In this paper's proposed system, a deleted reservation has a declared absence or retention disposition; a discarded deletion is a blocking discrepancy.
Transaction relationships matter when a business operation changes several records. A cancellation and its refund intent may be committed together in the source but appear temporarily separated in a candidate processing path. A comparison that catches the candidate halfway through application can falsely describe ordinary lag as a semantic defect. Conversely, a consumer that never applies the second change can look perpetually delayed. Require evidence of completion at the selected boundary, then classify the difference. Waiting indefinitely is not a recovery strategy.
Native replication also has scope limits. PostgreSQL's logical replication restrictions describe schema, sequence and large-object limitations. A copied identity column does not establish that a future writer has a safe sequence state. Treat schema readiness and writer initialization as separate acceptance items. Provider-supported replication should be used where appropriate, but its support for copying particular data is not a certification that the replacement application can safely accept every business operation.
Use provider validation inside its documented scope
Provider validation can reduce bespoke work when supported endpoint, key and data-type conditions fit the migration. AWS DMS describes source-to-target row comparisons and reports validation outcomes, with additional resource consumption and documented limitations. Its data validation guide is a useful starting point for verifying a copied table. It does not define the hypothetical booking platform's cross-entity capacity or refund contract. Keep that semantic review independent.
Record the actual validation configuration in the evidence packet. A passing setting that compares only part of a large value cannot establish agreement for the omitted remainder. AWS documents partial LOB validation and validation-query delay options in validation task settings. Use those settings deliberately and describe their limits. Pending or suspended checks are not equivalent to validated records. Configuration choices should remain visible beside the result rather than buried in a task console.
Preserve recovery metadata before operational cleanup. AWS DMS ongoing replication guidance describes native start points and recovery checkpoints, including metadata lost when a task is deleted. A replication checkpoint is evidence about that task's stream recovery point. It is not, by itself, proof that application invariants or external payment effects are reconciled. The proposed comparison binds task evidence to separately observed business state and retains both under controlled access.
Budget validation work against live-system capacity. Read-heavy checks can compete with customer requests and add network or storage pressure. Set execution windows, resource limits and stop conditions using the actual service contract. If the team pauses comparison to protect production, record the incomplete population and evidence age. Do not quietly reduce sampling or omit slow partitions while leaving the report labelled complete. A slower honest review is preferable to a fast report that changed its meaning without approval.
Classify divergence before selecting a repair
Use an explicit discrepancy taxonomy. Incomparable observations require another defensible pair, not a data patch. Capture lag requires continuity and applied-state evidence. A representation difference needs review against the declared mapping. A business-invariant violation needs domain analysis. Missing evidence remains unknown until the responsible owner can establish what occurred. Separating these classes prevents a migration operator from repairing a record whose only problem was that the two observations were taken at different states.
Keep each discrepancy tied to the tenant, business identity, manifests, mapping revision and invariant. Record source and candidate evidence references, not unnecessary raw personal information inside tickets. Assign an owner, next permitted action and resolution condition. The closure should explain whether the record was corrected, the mapping changed, the population was legitimately excluded or the comparison was replaced. An operator's comment that everything looks fine is not reproducible evidence for the next reviewer.
Avoid automatic source-wins repair when the business model intentionally changes. Source data may be malformed, the target may contain a legitimate post-cutover write, or the observed difference may concern an external effect rather than a stored field. Repair is a state-changing operation with its own scope, authority and concurrency conditions. Use a conditional repair against the observed version where appropriate, then re-extract and re-evaluate affected invariants. A successful update count proves a database transition, not that the business contract is now satisfied.
Set a bounded recovery path. A repeatedly failing partition should remain held with an owned investigation rather than retrying at full load indefinitely. Escalate when history is missing, authority is ambiguous, the mapping is unstable or an effect cannot be settled. The paper's proposed state flow intentionally has no automatic path from discrepancy to business write authority. Detection, repair authorization, refreshed evidence and cutover acceptance are separate decisions, even if one person participates in several of them.
Keep evidence recovery separate from write authority
The second visual is a decision lifecycle, not another system topology. It follows an observation from pairing through comparison, then branches between scoped acceptance and owned divergence. A repaired discrepancy returns through a fresh paired observation. Acceptance produces a review record, while the existing writer remains unchanged until a separate cutover gate checks current conditions. This prevents an old passing report from becoming implicit authorization for a later transfer.
Record what invalidates acceptance. A new schema, changed tolerance, delayed deletion, additional writer or altered downstream adapter can change the conclusion. The gate should verify the reviewed scope and evidence still apply to the current candidate. Do not substitute a label such as validation complete for the specific revisions and cut that passed. If conditions changed, the right action is to repeat the affected review, not extend the old evidence to an unexamined state.
For an unknown external effect, the lifecycle stops short of another attempt. A refund request that timed out might already exist at the payment destination. Preserve its operation identity and use supported destination readback or an owned settlement process. Replaying the booking change under a new refund identifier can duplicate an effect while making the local records appear cleaner. Unknown is a state to resolve, not a synonym for failed. Keep that distinction in both the discrepancy queue and the customer-facing status.
Prove the comparison with independent fixtures
Start with small, adversarial synthetic histories rather than only a large successful dataset. Include a cancelled booking, a delayed refund intent, a deleted reservation, a tenant-qualified identity collision and a schema change that makes a formerly permitted enum unknown. Use inert adapters so the exercise cannot produce real refunds or notifications. The expected result should identify which business invariant passes, which discrepancy blocks acceptance and what evidence is still absent.
Deliberately test a bad comparator. If a mapper drops cancelled bookings, a row-count-only comparison or shared projection bug may hide the defect. A fixture should fail that implementation. Include a case where both implementations agree on an incorrect capacity outcome to show why the independent invariant matters. This is not an assertion that the proposed comparator is correct. It is a way to make its intended rejection behavior explicit before relying on it.
| Synthetic condition | Required disposition | Evidence to retain | | --- | --- | --- | | Candidate cut omits a committed cancellation | Inconclusive pair; no cutover acceptance | Source cut, candidate applied state and fresh pairing method | | Both systems show a cancelled booking as active | Fail the independent capacity invariant | Accepted cancellation policy and known input history | | Candidate loses a deletion after transformation | Blocking lifecycle discrepancy | Delete lineage, transformation revision and retained-state rule | | Same booking key appears in two tenants | Compare tenant-qualified identities separately | Tenant scope and each reservation's evidence | | Refund attempt has no conclusive destination receipt | Hold the effect as unknown; do not mint another operation | Original operation identity and owned readback path |
Retain every intended fixture in the denominator. Distinguish executed, passed, failed, inconclusive and not observed cases. When the denominator is zero, there is no agreement rate. For a sampled production-shaped exercise, report coverage by tenant, lifecycle and state transition rather than implying untested history passed. A complete-looking percentage can hide the exact cancellation or settlement path that makes the migration unsafe. The decision needs evidence coverage and correctness, not a number detached from its population.
Fence the cutover and define the recovery limit
Before transfer, identify every path that can write scoped state: APIs, scheduled tasks, internal tools, partner callbacks and repair jobs. A routing switch changes where requests go but does not necessarily stop an old worker. The proposed contract requires one authorized writer per partition and an observed fence against stale writer activity. Verify the fence under concurrent and delayed requests. A control-plane configuration saying the old application is disabled is weaker evidence than a tested rejection at the actual write boundary.
Settle or account for in-flight work before accepting the final cut. An old request may commit after the user receives a timeout. A scheduled job may have claimed work before the fence. A payment callback may arrive after the application switch. Each needs an operation identity and a permitted disposition under the new authority model. Data agreement measured before these effects settle cannot be presented as proof of their eventual outcome. The cutover record should name any accepted pending work and who owns it.
After the first new authoritative write, rollback is not merely directing traffic to the old application. The old state may not contain that write, may not represent the new schema or may replay an external effect incorrectly. Define whether recovery uses reverse transformation, forward repair, a temporary hold or an approved loss boundary. Rehearse the selected path using synthetic writes. If the team cannot explain how a new reservation survives recovery, the proposal does not yet have a defensible rollback contract.
Avoid two-way synchronization as a casual escape hatch. Conflicting writers need explicit ownership, version and conflict rules. A last-write-wins policy based on timestamps can overwrite a valid cancellation or money state if those timestamps are not comparable business authority. If active dual writing is a real requirement, review it as a separate architecture decision. This paper's one-writer assumption does not apply to an unexplained multi-writer system, and the diagrams should not be read as making that system safe.
Security and operational evidence handling
The reconciliation service can expose more data than an ordinary application request because it reads across many tenants and histories. Scope its identities, stores and reports to the approved population. Separate read access from repair and cutover permissions. Protect evidence at rest and in transit using the organization's chosen controls, and record access to sensitive extracts. A debugging export should not become an ungoverned second copy of customer information on an engineer's laptop.
Retention needs a purpose and an end. Keep enough evidence to reproduce the accepted decision and investigate an approved discrepancy, but do not preserve personal data indefinitely simply because migration might someday need it. Prefer references and minimized projections when full payloads are unnecessary. Explain how deletion obligations reach comparison stores, report attachments and temporary artifacts. Removing a record from the replacement does not remove copies already exported for reconciliation.
AI can assist with drafting discrepancy summaries or grouping already observed patterns, but it should not invent missing records, authorize a repair or approve cutover. Any generated explanation must point to the exact manifests and invariant evidence. Sensitive payloads should not be sent to an external model without an approved data-handling basis. A fluent explanation of why a difference is probably harmless is not evidence that the condition satisfies the business contract. The deterministic comparison and responsible reviewer remain the decision boundary.
Operationally, publish coverage, correctness, freshness and effect settlement as distinct dimensions. Track unresolved discrepancies by consequence and age, with named owners. Record interrupted runs and resource-driven pauses without turning them into passes. The useful business result is a defensible migration decision, not the maximum number of compared rows per second. This approach does not guarantee a lossless migration; it makes the conditions, evidence and remaining risk explicit enough for the accountable team to decide.
Review checklist and next useful step
Use this review checklist before accepting the comparison: is the scope versioned, are identities tenant-safe, are expected invariants independent of candidate outputs, do both manifests represent comparable committed states, are rejected observations included, and are mappings and tolerances fixed? Confirm deletes, schema changes, transaction completion and required non-database assets have declared treatments. Keep excluded populations visible with an owner and reason. Check that the conclusion remains limited to what was actually observed.
Then review the consequence of moving authority. Are every writer and in-flight operation accounted for? Does the write fence reject stale activity? Are unresolved external effects held under their original identities? Can recovery preserve writes accepted by the candidate? Has the team rehearsed a deliberately failing comparison and the proposed recovery path using inert adapters? If an answer is unknown, leave the acceptance gate open. A review checklist is useful only when missing evidence can stop the decision.
The next useful artifact is one completed comparison packet for a narrow business lifecycle, such as a booking that is created, cancelled and followed by a refund intent. Include its accepted invariant, independently known input history, paired manifests, mapping revision, discrepancy disposition and recovery boundary. Expand scope only after that packet is reproducible. Bring the packet and unanswered questions to an architecture review, rather than bringing only a migration dashboard or a diagram of connected databases.
For the owned execution procedure, use the dual-run reconciliation playbook. For the distinction between agreement and authority, read the old and new system comparison discussion. Ampity's platform modernization service and system architecture design service provide relevant routes for discussing scope and accepted delivery. Reading and PDF download require no contact details; an enquiry is optional and does not establish repair or cutover authorization.