Compare Old and New Systems Before Changing Authority

Compare replacement-system decisions at matched input revisions, classify differences and preserve missing evidence before approving a migration.

Choose the business result that must remain correct

A replacement service can return the same number of records as the old service while quoting a different price, admitting the wrong customer or dropping an approved cancellation. Compare the decisions that people rely on, using identified inputs and an explicit rule for acceptable differences. Keep the old service authoritative until the migration owner accepts the replacement's evidence at the intended scope.

Consider a hypothetical supplier-order service moving from a legacy application to a replacement. Both calculate delivery eligibility and a final amount. The replacement has a different data model and may use an AI assistant to extract supplier instructions. A row-count comparison cannot establish whether either service applied the right instruction, price revision or permission. The domain owner needs to identify the decisions that must agree and any deliberate behavior changes.

This article proposes a comparison contract for that review. The examples and counts are synthetic, not Ampity customer results or measured product performance. It concentrates on application outputs and decision meaning. Database copying, write transfer and post-cutover recovery require their own procedures. A successful comparison run can support one migration decision without proving every path is ready.

Pair inputs before comparing outputs

Give each comparison a stable identifier and record the input revision, actor, tenant, applicable policy and relevant dependency state. The two services should evaluate equivalent evidence. If the legacy service sees price revision eight while the replacement sees revision nine, their amounts cannot establish agreement or disagreement under one pricing rule.

Use a reproducible fixture where possible. Capture the permitted input and a controlled dependency response, then send that evidence through both calculation paths. For live comparisons, identify which changing fields are observable and which remain uncertain. Record the point at which the input was captured. A timestamp alone may not identify an inventory revision or prove that both reads saw it.

Treat missing pairing evidence as a separate result. A candidate response without its input revision remains unpaired, even if its amount happens to match. Excluding unpaired requests from the denominator can make coverage appear stronger than it is. Record captured requests, successfully paired requests, completed comparisons and untested paths separately.

AWS DMS's validation documentation describes corresponding-row comparisons and separate pending, failed and suspended states. It notes that continuously changing records can prevent comparison, and that validation adds database and network work. For application-output review, our recommendation is to preserve a similarly explicit unknown state, while defining the additional domain checks that a row comparison does not cover. Verify the tool's support for the actual endpoint and configuration rather than assuming one validation mode covers every migration.

Define permitted transformations field by field

Write the comparison rule before inspecting a large mismatch report. Identify fields that must match exactly, fields that need a reviewed conversion and fields that may legitimately differ. Preserve the original values alongside the normalized representation so a reviewer can explain what the comparison changed.

For an amount, include currency, unit, tax basis, rounding rule and the point at which rounding occurs. Converting dollars to minor units may be valid under a known contract. Rounding two different amounts until they match can hide a pricing defect. A tolerance needs a business reason, an owner and a defined scope; it should not be chosen because it reduces the current mismatch count.

For timestamps, distinguish an event's effective time from ingestion and processing times. A replacement can process the same instruction later without changing its validity. Conversely, an expired authorization should not become valid because the new service uses processing time. Compare the transition rule as well as the formatted output.

Ordering also needs a rule. Sort a list only when its order has no business meaning. A ranked eligibility list, prioritized fulfillment queue or sequence of revisions can change behavior when reordered. Treat missing, null, empty and redacted values separately where the contract does. Avoid a generic cleanup function that quietly collapses these states.

Separate calculation from consequential execution

Give the comparison path only the authority it needs. A shadow order calculation must not charge a customer, notify a supplier or reserve stock merely because the production path does. Inventory every external action and test how the comparison environment prevents it. An application flag is insufficient evidence if a forgotten worker still has production credentials.

Envoy's request-mirroring API documentation describes sending shadow traffic without waiting for its response before returning the primary response. Our recommendation is to instrument paired outputs explicitly. Receiving mirrored requests or seeing a healthy shadow cluster does not establish that a comparison completed, and the mirror itself does not grant safe side-effect isolation. Check the deployed Envoy version and route configuration before depending on the documented behavior.

Use restricted credentials, approved test adapters and destination controls for a consequential candidate path. Keep production data protection in scope even when writes are disabled. A mirrored payload can still disclose personal information or consume shared dependency capacity. Bound the comparison rate, diagnostic retention and access to discrepancy records.

Test the stop mechanism independently. If a candidate creates an unexpected effect or shared dependency pressure rises, stop further admission and record already admitted work. Disabling a traffic copy does not establish that a pending downstream request was cancelled. Reconcile its actual result before rerunning the same fixture.

Classify differences by consequence

The following worksheet is a proposed application-review artifact. It does not define automatic vendor behavior or replace the domain owner's acceptance decision. Keep classifications tied to a rule revision so a later rule change does not silently reclassify earlier evidence.

| Observed difference | Review disposition | Evidence before acceptance | | --- | --- | --- | | Formatting differs; meaning is unchanged | Apply the approved canonical conversion | Original values, conversion rule and tested examples | | Candidate implements a deliberate rule change | Review as changed behavior | Domain-owner approval and compatibility evidence | | Business invariant fails | Hold the affected capability | Reproducible fixture, correction and regression result | | Inputs or dependency revisions do not match | Mark the result unpaired | A matched comparison or explicit coverage limitation | | Candidate response is absent or incomplete | Keep comparison unresolved | Execution evidence, failure cause and a bounded rerun |

Do not assume the legacy result is always correct. If it violates an agreed rule, record that defect and decide whether the replacement should preserve compatibility temporarily or implement a separately approved correction. A modernization review cannot silently change a customer's entitlement because the candidate appears more sensible.

Inspect severity by consequence, not just difference size. A one-cent rounding discrepancy and a permission failure need different owners and release decisions. An aggregate agreement percentage should retain those categories. A small number of access violations can prevent acceptance even when most benign calculations agree.

Work through a bounded comparison ledger

Suppose an isolated fixture contains 200 identified quote requests intended for paired comparison. The observer retains a ledger entry for every request, including gaps in input or response evidence. The review records 160 exact agreements, 18 formatting differences covered by a previously approved conversion, eight deliberate rule changes, six invariant failures, five unpaired results and three missing candidate responses. These disjoint groups total 200.

For this fixture, the 160 exact agreements and 18 valid conversions provide 178 agreeing results within the existing rule contract. The eight changed behaviors still require their own acceptance. The six failures, five unpaired results and three missing responses remain separate open categories. Reporting 178 out of 178 by excluding every other category would hide 22 requests from the original scope.

Do not automatically label the remaining 22 as candidate defects either. Five lack the evidence needed for comparison, and three have an execution or observation gap. Investigate each cause, correct the fixture or implementation where necessary and preserve the original run. When a rerun is accepted, link its evidence to the affected comparison identities rather than adding repeated attempts to the denominator.

Add a targeted fixture for each confirmed failure. If a cancelled supplier instruction was accepted, test the cancellation revision, its effective time and the application identity that tried to act. If a price differs, retain the source line items and calculation basis. A generic test asserting that the next response matches will not necessarily catch the same defect after another rule change.

Use AI to investigate differences without approving them

An AI assistant can group recurring discrepancy descriptions or propose an explanation from permitted diagnostic records. Keep its suggestion separate from the measured classification. Require a source reference for any claim about the applicable business rule, and verify suggested normalizations before placing them in the comparator.

For a replacement that uses model output, define acceptance at the task level. An extracted supplier instruction needs the right fields, source support, permission and resulting deterministic checks. Two fluent summaries can disagree about the requested delivery date. Conversely, differently worded summaries may support the same authorized calculation. Compare the evidence and allowed result, not a superficial similarity score alone.

Keep repeated model runs distinguishable. Record the model and prompt configuration, retrieved input and attempt identity permitted by the diagnostic policy. If a candidate sometimes violates an invariant, reporting only its best attempt misrepresents the behavior. Review uncertainty and failure distribution under the declared test scope, without inventing a universal pass percentage.

AI review also has a privacy boundary. Do not send raw customer payloads to a new provider merely to explain a mismatch. Use approved data handling, restrict sensitive fields and preserve the human owner's authority over changed business rules. If the assistant or its evidence source is unavailable, leave the discrepancy open and continue the deterministic review.

Accept only the capabilities the evidence covers

Sampling can find defects but cannot prove that untested tenants, roles, seasonal workflows or rare transitions agree. Record the covered cohorts and intentionally absent ones. Exercise a denied read, expired instruction, corrected price, cancellation and missing dependency response before asking the owner to accept the supplier-order capability.

Agree stop conditions for the review, including unauthorized effects, lost pairing evidence and shared-service pressure. Separate comparator defects from candidate defects and correct either before expanding the sample. Do not loosen a tolerance or omit a failing field during a run without preserving the old rule and obtaining approval for the revised one.

Start the next review with one capability and a versioned comparison contract: input pairing, field rules, invariants, open classifications, observed coverage and decision owner. Use the database migration guide for data-copy and authority-transfer boundaries, and the capability migration playbook for phased replacement. Bring one completed discrepancy ledger and the unresolved questions to a platform modernization review or system architecture review. Reading these resources does not require contact information.