Rehearse a DMS Stop and Resume Without Losing the Recovery Evidence
Establish a full-load baseline, deliberately stop a preserved full-load-and-CDC task, and reconcile known source changes before calling the rehearsal complete.
To rehearse DMS recovery, preserve one previously executed full-load-and-CDC task, establish a reconciled full-load baseline, stop its CDC phase under separate authorization, and compare an independently expected insert/update/delete workload after resumption. A task returning to running is a control-plane observation, not evidence that the target is complete or correct. Keep the target non-authoritative throughout. Missing source logs, an unexplained checkpoint, changed task identity or unresolved data differences hold acceptance.
This procedure concerns task-based AWS DMS on a replication instance, from one Amazon RDS for MySQL 8.0 instance writer to a separate RDS for MySQL 8.0 writable instance. Actual complete database versions and DMS engine version must be collected before execution. The tables are ordinary InnoDB tables with integer primary keys, no LOBs, no transformations and a frozen schema. Aurora, read-replica sources, multi-writer designs, Serverless, homogeneous migrations, DDL changes and application cutover are outside this run. Every worked identity and result below is fictional. No AWS action, database statement or outage was executed for this paper.
1. Choose the interruption you can actually observe
Start by naming the proposed failure claim precisely. This run tests a planned stop after full load and cached-change application, followed by same-task resumption while source changes were committed during the stop. It does not test a process crash, lost replication storage, unavailable source, failover, task recreation or a stop halfway through copying a table. Calling all of those “restart” hides materially different recovery conditions. Record excluded modes in the packet so a later migration review cannot turn one favorable result into broad outage assurance.
Choose an isolated rehearsal environment that reproduces relevant schema, endpoint settings and access boundaries without carrying private production payloads. Use generated records and disable target-side notifications, payments and integrations. The source remains the sole application writer. The target may receive DMS writes but never business writes. An application owner verifies this boundary, rather than relying on a database name containing “test.” If isolation cannot be demonstrated, stop before scheduling the interruption.
Capture complete source and target engine strings, DMS engine version, Region, task ARN alias, replication-instance identity, endpoint aliases, mappings and task settings. Family-level endpoint support does not establish patch availability, connectivity or compatibility for the actual proposal. An UNKNOWN version holds execution. Do not select an arbitrary documented release merely to fill the field.
2. Assign separate operational and acceptance authority
The rehearsal lead approves the bounded plan with the environment owners. The DMS operator receives permission for the named task actions only. The source DBA owns test mutations and source log evidence; the target DBA owns read-only comparisons and recovery of the disposable target. The application owner defines the invariant before observing results. An independent observer records whether the actual sequence matched the plan and whether the packet is complete. One person can hold several roles in a small team, but the author of a favorable result should not quietly become its sole acceptor.
The action distinction matters. StartReplicationTask distinguishes first execution, resumption and reload. This run starts the new full-load-and-CDC task with start-replication, then uses resume-processing only after that same task has executed and stopped. reload-target is not an interchangeable recovery button: it reloads all tables and starts capturing changes. Any proposed reload, task replacement, table repair or new checkpoint start requires a separate decision, not an improvised retry under the resume approval.
Agree a maximum interruption window, stop-request timeout, resume-request timeout, monitoring interval, target catch-up deadline and source storage threshold. These are supplied operational limits, not universal AWS defaults. Record who can halt generated writes and who can restore the isolated environment. A missed observation deadline changes the result to inconclusive even if the target looks correct the following morning. Otherwise the rehearsal proves no useful recovery-time boundary.
3. Close permissions, log coverage and target prerequisites
Have the source DBA review the current MySQL source prerequisites for the observed deployment. Retain evidence of required backup/binlog configuration, row-based full images, endpoint access and relevant existing-session behavior. Do not copy broad privilege grants from this playbook, change production parameters or restart sessions without approval. A denied read is a missing observation, not proof that the configuration is safe.
The target DBA separately reviews MySQL target prerequisites, including load access, local_infile and full-load foreign-key handling. Any temporary control change needs a recorded restoration owner and acceptance check. This run uses ordinary pre-agreed tables and does not claim that DMS recreates the complete application schema. Preserve indexes, constraints, defaults and triggers as explicit schema evidence, not assumptions derived from a successful table copy.
Use the retention-budget decision from the DMS source-log retention analysis, but require current evidence of the actually available log range before interruption and before resumption. A configured retention duration is not proof that the needed recovery position is readable. The DBA records how the saved task checkpoint was related to available source history and any uncertainty. If the team cannot substantiate that relation, hold rather than inventing an equality between a DMS token and a wall-clock timestamp. Recheck available storage while the task is stopped. The RDS log configuration is a policy input, not a substitute for those observations.
For a same-engine migration, also ask why task-based DMS is the selected approach. The source documentation points readers toward homogeneous migration tooling for MySQL-to-MySQL scenarios. Native replication, snapshot-based transfer or homogeneous migration may be better alternatives depending on constraints. This playbook tests an existing task-based choice; it does not establish that choice as the best architecture.
4. Establish the full-load baseline before introducing the stop
Preserve a versioned export of mappings, endpoint attributes, schema and task settings before first start. Record target-table preparation mode and both cached-change stop settings. Full-load settings distinguish stopping before cached changes from stopping after them. This rehearsal does not silently change either flag to manufacture a baseline. The operator documents the chosen configuration and its expected transition under the current task documentation.
The fictional J1 plan uses DO_NOTHING with separately prepared, empty target tables and both StopTaskCachedChangesApplied and StopTaskCachedChangesNotApplied false. This permits the planned run to continue into CDC before the operator's stop request. It is not a default recommendation for an existing populated target. If actual settings contain an automatic stop, nonempty target or different preparation mode, the owner must revise the baseline and action plan before starting. Never infer safe target contents from a mode name.
After initial start, collect every selected table's statistics, including all pagination from DescribeTableStatistics. Check for queued, loading, suspended or errored populations. “Full-load 100%” alone does not prove that cached changes have been applied or that the complete mapped population is correct. Record task observations, per-table observations and the independent baseline comparison as distinct evidence files with timestamps and collection scope.
Pause the controlled source writer at a declared committed cut for baseline comparison. Confirm no open generated transaction can later change that cut, and wait for the agreed target observation conditions. Export complete selected rows from both sides using the same projection and deterministic key ordering. The observer checks full population coverage and the business invariant below. If data continues changing or the cut is ambiguous, the baseline is unpaired and the interruption must not proceed. Do not add arbitrary sleep time and label the result synchronized.
5. Predeclare the workload and expected results
The fictional baseline has two tables. orders contains O101 with total 100 and O102 with total 50. order_lines contains L1/O101/100 and L2/O102/50. These identifiers are aliases for generated integer-key rows, not proposed database types. The invariant is that each remaining order total equals the sum of its lines, and no line refers to a missing order. Both tables must participate in comparisons. Equal total row counts alone cannot distinguish an update from a missed delete and an extra insert.
During the stop, transaction T1 updates O101 and L1 from 100 to 120. T2 inserts O103/75 and L3/O103/75. T3 deletes L2 and O102. Each transaction is independently committed, with source commit evidence and an immutable operation identifier in the test ledger. The expected final orders are O101/120 and O103/75; expected lines are L1/O101/120 and L3/O103/75. Each table has two rows and total value 195, up from 150. There are no O102 or L2 rows. These expectations are stipulated before the run, not calculated from a favorable target export afterward.
Arrows show the proposed evidence sequence, not an observed AWS run. The lower decision card includes HOLD when evidence is missing, failed or ambiguous; that stop condition applies at every stage. Running is separated from reconciled, and neither state grants target authority.
Use a bounded transaction generator approved by the source owner. It must not retry a timed-out write blindly: first resolve whether that transaction committed. A client error is not proof of rollback. If commit status is unknown, suspend new writes and mark the packet ambiguous. This specimen intentionally omits concurrent writers, LOB truncation, large transactions and schema changes. A production workload with those characteristics needs additional separately expected cases; the simple ledger is not representative-performance evidence.
6. Stop the preserved task and record the gap
At the approved boundary, the operator requests the named task stop and preserves the request identity and response. The StopReplicationTask example can return a stopping status, so the observer waits for a stable stopped readback under the agreed timeout. A failed request or a task that never reaches the required state ends this planned variant. Do not proceed with the mutations as though an interruption had been established.
Once stopped, save the complete task readback, stop reason, failure message if present, recovery checkpoint, settings and mappings. The checkpoint is opaque recovery evidence. AWS documents that the latest checkpoint can be obtained for a stopped or failed task, and that deleting the task loses checkpoint information. Preserve the task and its evidence. Do not delete and recreate it, rewrite the token or replace it with a guessed source position.
The source owner now performs T1, T2 and T3 within the approved window. The observer records start and finish times, commit receipts and any retry ambiguity. Pause the controlled writer after T3 at cut C1; preserve a source export tied to that cut. Verify task identity has not changed while another administrator was working. Check source storage and retained history before consuming the remainder of the interruption budget. If a storage alarm requires ending the run, halt generated writes and use the agreed recovery decision, not an unreviewed retention extension.
7. Resume once, then observe without repairing the evidence
Before resume, compare task, endpoint, schema, mapping and settings identities to the baseline. An unexplained change invalidates attribution. Confirm the saved source history remains available and endpoint connectivity/access evidence is current. The operator requests resume-processing for the same previously executed task and saves the response. Observe subsequent status, logs and table evidence through the monitoring window. A starting or running response is not the end of the procedure.
Do not issue repeated starts because the target has not changed immediately. First determine the current task state and the first request's outcome. Record source-capture and target-apply observations separately, including unresolved errors and suspended tables. Lag metrics can help diagnose progress, but they do not establish the saved checkpoint's log coverage or prove the business cut was applied. If the declared catch-up deadline expires, preserve the partial target and classify the run as failed or inconclusive according to the evidence, rather than changing the deadline after seeing the result.
The target remains read-only to applications while DMS writes. Do not manually insert a missing row, erase a duplicate or repair a value before preserving the divergence. A repair would replace the outcome being tested. If the recovery owner later authorizes repair or reload, create a new run identity with new baseline and expectations. The old failed run remains in the packet.
8. Compare the paired cut, not two convenient screenshots
With the controlled source writer still paused at C1, obtain the target export after the declared observation conditions. Bind both exports to run identity, schema, mapping, projection, table/key population, collection times and query definitions. The source export alone does not prove the target's relationship to C1; the packet needs the applied-progress observations and the complete expected-key comparison. Use exact generated values, deterministic ordering and explicit missing/extra-row results. If the exports are truncated, permission-filtered or collected against different populations, hold.
Check T1's two values, T2's two insertions and T3's two absences individually. Then check each order-line invariant and both table totals. A target showing two orders totaling 195 can still be wrong if it contains O102/75 instead of O103/75. A missing delete can hide behind an unrelated missing insert. Preserve per-key evidence rather than relying on a single checksum whose canonicalization, coverage or collision assumptions were never agreed.
The supplied offline specimen verifies only the fictional transaction ledger and its expected snapshots. It cannot query a database, parse a DMS checkpoint, prove a transaction committed, establish log availability or authorize resumption. The team must attach actual observations to the editable packet. Passing the specimen is useful protection against a mistaken expected result, not a migration test result.
9. Keep DMS validation and business acceptance separate
Where enabled and supported, collect table-level DMS validation evidence alongside the independent checks. Validation adds workload and requires suitable keys. Disabled, pending, suspended, mismatched or table-error evidence is not a pass. Its state can change when rows change; retain the collection time and table scope. External target writes also compromise the comparison boundary.
A favorable row comparison does not establish the order-line business invariant, absence of target effects or migration authority. Conversely, a correct synthetic invariant does not erase provider-reported validation failures. Resolve discrepancies between the two evidence surfaces explicitly. An excluded table or field needs an owned alternative check and approved scope limitation, not a blank cell interpreted as passed.
Do not enable automatic repair merely to turn the dashboard green during this rehearsal. Any correction feature changes the intervention being assessed and needs a separately specified variant. Preserve original differences and demonstrate the corrected packet at a new paired cut if a repair is authorized later.
10. Decide failure, inconclusive or bounded acceptance
If required source logs are missing, stop repeated resume attempts and preserve task/checkpoint/source evidence. The source remains authoritative. A clean reseed may be a legitimate recovery alternative, but it requires its own disposable-target, table-preparation and validation decision. It does not retroactively prove same-task continuity. Likewise, task deletion, changed mapping, unknown commit status and an unpaired cut make this run inconclusive even if selected target rows happen to match.
If the task resumes but a delete is absent, record a failed reconciliation with exact keys, expected values, actual values and scope. If validation cannot finish within the resource budget, retain an incomplete result rather than disabling validation and approving the run. If target effects occurred, involve the application owner immediately; simple target disposal may not reverse external actions. That boundary is why effect isolation was a prerequisite.
A bounded accepted result means only that this exact planned stop/resume sequence met its predeclared conditions in the observed environment. It is not proof of mid-full-load recovery, crash tolerance, every source transaction, peak throughput or application cutover readiness. Record those exclusions next to the conclusion, not only in an appendix that can be detached from the result.
11. Complete the owned decision record
Use the editable decision record to retain the chronology and evidence references. It contains both a complete blank record and a filled fictional run; the accompanying synthetic packet supplies the independently specified rows. The packet is an offline template, not recovery evidence. Evidence aliases refer to access-controlled artifacts whose contents, version and permission scope the team must verify. An unchanged alias is not an immutable guarantee; changed evidence invalidates the decision until reviewed.
Download the offline companion ZIP without providing contact details. It contains the editable record, synthetic packet, README and six local expectation tests. Extract it and run node --test expected.test.mjs with Node.js 18 or later. These tests use no AWS credentials or database connection and cannot establish an actual recovery result.
Filled example
- Run and accountable owner
- R7, fictional migration lead Mira; observer Dev. This is a proposed specimen, NOT EXECUTED.
- Environment and exact versions
- Separate RDS MySQL 8.0 instance writers; full engine strings and DMS engine UNKNOWN. Execution HOLD until collected.
- Task, schema and settings
- Task-A, schema-S1, mapping-M1, settings-J1 are synthetic aliases, unchanged across the proposed sequence.
- Baseline and authority
- C0 has O101/100, O102/50 and matching L1/L2; source sole writer, target effects disabled.
- Interruption and source evidence
- Planned CDC stop only; stable stopped status, checkpoint token and readable log coverage are required but NOT OBSERVED.
- Changes and paired result
- T1 update, T2 insert, T3 delete; expected C1 has O101/120 and O103/75, L1/120 and L3/75, no O102/L2.
- Acceptance and recovery
- Expected arithmetic 195 in each table. Actual outcome UNKNOWN; no operational acceptance. Preserve source authority and failed evidence; separately approve any reseed.
Blank record
- Run and accountable owner
- Record run identity, operator, independent observer, permissions and acceptance owner.
- Environment and exact versions
- Record complete source/target/DMS versions, Region, topology, isolation and exclusions; unknown required inputs hold execution.
- Task, schema and settings
- Record task and endpoint identities, schema/mapping/settings references, target preparation mode and change detection.
- Baseline and authority
- Record paired C0 exports, selected population, full-load/cached-change evidence and source-only application authority.
- Interruption and source evidence
- Record approved limits, stable stop, checkpoint, source log availability, storage observations and request/readback chronology.
- Changes and paired result
- Record predeclared transactions, source commit evidence, C1 cut, resume outcome, target export and per-key differences.
- Acceptance and recovery
- Record validation/business checks, disposition, excluded failure modes, restoration owner and separately authorized next action.
12. Close with an acceptance checklist and a specific next action
The acceptance owner checks that the packet contains the original plan, version/configuration evidence, complete baseline population, stable stop, preserved checkpoint, log-coverage attestation, known source commits, same-task resume observations, paired exports, per-key insert/update/delete results, business invariant and table validation scope. Every failed or unavailable item is explicitly failed or unknown. The observer confirms that no repair, target business write or unrecorded configuration change occurred between baseline and result.
Restore only controls whose reversal was pre-approved, such as the controlled writer pause or a temporary rehearsal-only access grant. Retain logs and evidence under the agreed retention and access policy before disposing of generated target data. Do not delete the task to tidy the console while the recovery investigation still needs its checkpoint. Cleanup is an owned action with its own authorization, not an automatic last step.
The next useful action is for the migration lead and source DBA to complete the version, permission and log-coverage fields, then ask the independent observer to challenge the predeclared expected results before scheduling the isolated run. If this planned interruption is accepted later, choose the next missing failure mode explicitly, such as a controlled source outage or an interrupted full-load variant. Do not expand the favorable claim without a new plan, independent expectations and observed evidence.