Will the Source Logs Outlast an AWS DMS Interruption?
Budget source-log age, capture recovery and storage for an interrupted AWS DMS task using RDS for MySQL. Keep replay evidence separate from target lag.
An AWS DMS restart is feasible only if the required source change range remains readable through the proposed recovery path. Budget the age of the oldest required position, the interruption, recovery work and an explicit reserve. Check the actual retained range as well as the configured retention period. Then check whether the database can afford the additional logs. A large retention value cannot repair a range that has already disappeared.
This article helps a database migration owner review that budget for a task-based AWS DMS migration with an RDS for MySQL 8.0 instance deployment as the source. The source remains the sole application writer; the target is a non-authoritative migration candidate. Exact deployed patch versions, DMS engine, endpoint support and permissions must be checked for the real migration. Aurora, MariaDB, PostgreSQL, MySQL 8.4, RDS Multi-AZ DB clusters, DMS Serverless and homogeneous data migrations are outside this worked scope. Their retention and recovery rules require separate review.
The numbers, resource aliases and log observations below are fictional. No database, DMS task or CloudWatch observation was created or queried. The arithmetic is an inspectable planning calculation, not an AWS performance benchmark, resume procedure or permission to change retention. Its output is a reviewed evidence request and a bounded interruption budget. Cutover still needs the separate reconciliation decision.
1. Establish which history the recovery needs
CDC reads changes from the source engine's logs. For the selected MySQL source, AWS documents row-based binary logging, full row images and the relevant CDC replication privileges. For RDS for MySQL, automated backups enable binary logging. Existing sessions can retain an earlier logging format after a parameter change, so a new parameter value alone does not establish coverage. Review the effective source behavior through the database owner's approved procedure. AWS DMS MySQL source requirements
The retention decision starts with a particular task and recovery position. Record their identities together with source instance identity, endpoint configuration revision, observation time and required transaction boundary. A convenient recent file is not a substitute for the position the recovery actually requires. If a source was restored, replaced or repointed, do not assume the same-looking file alias belongs to the same continuous history.
AWS describes a MySQL native CDC start position as a binary-log filename plus position. DMS also retains its own recovery checkpoint information, available through task readback; deletion of the task loses that task checkpoint evidence. Preserve the provider value without inventing a parser for its internal representation. A reviewer must establish how it relates to the required source range and the selected restart operation. CDC start points and checkpoints
Three different boundaries need separate records: the source history still available, the task's supported recovery point, and what the target has durably applied. A target may be behind changes already captured by DMS. Conversely, the capture path may be behind before those changes ever reach the target. A retention calculation cannot establish which boundary is limiting without evidence. Require the migration operator and database owner to agree which position the proposed recovery might reread, rather than estimating it from a target timestamp.
2. Read the configured window and available range separately
For RDS for MySQL, binlog retention hours defaults to NULL, which means zero hours of retained binary logs. AWS documents a maximum of 168 hours for MySQL DB instances and does not permit zero as the setting value. The configuration persists through reboot or failover, and AWS instructs operators to monitor the storage consumed by retention. Use the documented configuration readback through an authorized database identity. This article does not supply a state-changing call. RDS binary-log configuration
A configured 24-hour window is a policy input. An observed inventory proves a different fact: which required files and positions can be read now. Retain both. A new setting cannot demonstrate that older files still exist, and a file inventory without continuity and position evidence cannot demonstrate that the whole required interval exists. A retention change also needs a renewed capacity review rather than being treated as a harmless migration preference.
Use explicit evidence states. Available means the required position and continuing range have current supported read evidence. Missing means the needed range is demonstrably absent or unusable. Unknown means the observation cannot decide, for example because access was denied, the range was not fully listed, the checkpoint cannot be related to the source, or the source identity changed. Missing and unknown both hold a continuity claim, but they call for different next actions.
A backup success record does not establish that the active DMS source exposes the necessary binlog range. A restore may provide a different recovery option, with its own source identity, consistency and replay plan. Keep that option visible without silently substituting it for an ordinary task resume. Starting from a later position can omit changes; reloading can replace target state. Neither action belongs inside a spreadsheet's automatic recommendation.
3. Calculate a conservative fixed-position horizon
For an initial screening calculation, assume the oldest required replay position remains needed until capture catches up. This deliberately holds that position fixed instead of assuming it advances safely. Let A be its age now, D the total time from now until capture can make useful progress, C the subsequent capture catch-up duration and M the separately agreed reserve. All four quantities use hours. The screening retention requirement is A + D + C + M.
D includes detection, escalation, approval, repair, initialization and reconnect delays, not merely the planned maintenance interval. M covers a stated uncertainty, such as a second unsuccessful reconnect and the next decision meeting. Do not use the same incident allowance in both D and M without declaring the duplication. Nor should a round number conceal an unmeasured initialization or recovery phase.
This fixed-position horizon is a sufficient planning screen under its assumptions, not the minimum retention that every DMS task needs. A supported, evidenced recovery frontier may advance while capture progresses, reducing the oldest necessary source age. Demonstrating that requires a time series of the actual recovery relationship and retained range. It cannot be inferred from successful network traffic or target throughput. If that relationship is unknown, preserve the conservative assumption and label it as such.
Estimate catch-up only with comparable work units. In the synthetic model, source backlog is expressed in GiB of source-binlog work, while arrival and capture-processing rates use that same work definition in GiB/hour; one GiB is 1,073,741,824 bytes. These are stipulated accounting units for this example, not a conversion from DMS rows, network bandwidth or target storage. Assume ordered, complete work, constant rates and no additional transaction, filtering, compression or processing bottleneck.
Let B0 be the existing source-work backlog, r the continuing arrival rate and c the sustainable capture processing rate after recovery. Backlog at useful restart is B0 + r × D. If c is greater than r, catch-up duration is (B0 + r × D) ÷ (c − r). The denominator is spare capacity after new work. With a positive backlog at useful restart, c equal to or below r gives no finite catch-up time under these constant-rate assumptions. An evidenced zero backlog has no initial catch-up work; unknown backlog is not zero, and future arrivals still require adequate processing capacity. Returning to a running status cannot change that arithmetic.
Use B0 from bounded source-work evidence, not from the target's age. Only if a representative measurement supports a constant rate over the entire relevant age interval can B0 be approximated by r × A. A large transaction or rate spike can invalidate that approximation. Retain its method, included workload, time window and uncertainty; unknown backlog or incompatible units means unknown C, not zero hours.
4. Work through an interruption and its opposing cases
Suppose fictional source SRC-17 has a required replay point two hours old. The example stipulates a 24 GiB existing backlog, 12 GiB/hour continuing arrivals and 30 GiB/hour sustainable post-recovery capture processing. Useful capture will resume six hours from now. The owner adds four hours of reserve and proposes a 24-hour retention window. A fresh fictional inventory is assumed to cover the required range continuously. None of those inputs is an observed AWS result.
The six-hour interruption adds 72 GiB, so restart backlog is 96 GiB. Spare processing is 18 GiB/hour: 30 minus 12. Catch-up takes 96 ÷ 18, or 5 hours 20 minutes. While catching up, another 64 GiB arrives. Total processed is therefore 160 GiB, equal to the restart backlog of 96 plus those 64 new GiB. This reconciliation checks the same-unit model rather than relying on a rounded duration.
The fixed-position retention screen is 2 + 6 + 5⅓ + 4 = 17 hours 20 minutes. A 24-hour policy leaves 6 hours 40 minutes beyond that screened horizon. This is conditional arithmetic only. It does not establish log readability, supported task recovery, target correctness or approved storage capacity.
| Independently stipulated case | Calculated consequence | Review disposition |
|---|---|---|
| Base: B0 24 GiB, r 12, c 30, D 6, A 2, M 4 hours | Restart 96 GiB; catch-up 5 hours 20 minutes; horizon 17 hours 20 minutes | Time screen fits 24 hours; range and storage gates remain separate |
| Higher change: B0 40 GiB, r 20, c 30; other times unchanged | Restart 160 GiB; catch-up 16 hours; horizon 28 hours | 24-hour fixed-position screen fails by 4 hours |
| No spare capture: c 12 equals r 12 | Restart 96 GiB; backlog cannot shrink under constant rates | No finite catch-up estimate; HOLD the proposed plan |
| Required file absent despite a 24-hour setting | Arithmetic remains unchanged but continuity evidence fails | HOLD resume continuity; investigate an authorized recovery or new baseline |
| Task captured ahead but target apply is slow | Target backlog alone does not specify source replay age | Obtain actual recovery-position and source-range evidence |
- Base: B0 24 GiB, r 12, c 30, D 6, A 2, M 4 hours
- Calculated consequence: Restart 96 GiB; catch-up 5 hours 20 minutes; horizon 17 hours 20 minutes
- Review disposition: Time screen fits 24 hours; range and storage gates remain separate
- Higher change: B0 40 GiB, r 20, c 30; other times unchanged
- Calculated consequence: Restart 160 GiB; catch-up 16 hours; horizon 28 hours
- Review disposition: 24-hour fixed-position screen fails by 4 hours
- No spare capture: c 12 equals r 12
- Calculated consequence: Restart 96 GiB; backlog cannot shrink under constant rates
- Review disposition: No finite catch-up estimate; HOLD the proposed plan
- Required file absent despite a 24-hour setting
- Calculated consequence: Arithmetic remains unchanged but continuity evidence fails
- Review disposition: HOLD resume continuity; investigate an authorized recovery or new baseline
- Task captured ahead but target apply is slow
- Calculated consequence: Target backlog alone does not specify source replay age
- Review disposition: Obtain actual recovery-position and source-range evidence
The higher-change case deliberately changes the starting backlog as well as arrivals: it stipulates two hours at 20 GiB/hour. Reusing the base case's 24 GiB without stating that different history would muddle the comparison. A second alternative could retain B0 at 24 and change only future arrivals; record it as a different scenario rather than calling both scenarios peak load.
5. Reject a time budget that the source cannot store
Retention trades recovery room against source storage and operating risk. In the base synthetic constant-rate model, a 24-hour retained window contains approximately 12 × 24 = 288 GiB of binlog work. If the approved budget for retained logs is 200 GiB, the proposed window fails the storage screen by 88 GiB, despite passing the time screen. That is the disposition of this complete worked case: HOLD the combined proposal.
The 200 GiB allowance is a fictional dedicated budget, not total free disk or an AWS threshold. Real headroom must also cover database growth, temporary work, backups where applicable, existing files, burst reserve and operating alarms. The model's source-work units are stipulated to match retained physical log bytes only for this storage example. In a real assessment, use actual binlog generation and storage evidence instead of assuming capture work and physical storage have a universal conversion.
At the higher rate, 24 hours implies 480 GiB under the same simplified storage assumption. A 28-hour screened horizon would imply 560 GiB before any additional allowance. Longer retention alone therefore does not resolve that opposing case. The team might reduce interruption time, improve demonstrated capture capacity, provide approved storage headroom, or select a different migration and baseline strategy. Each alternative needs its own evidence and source-impact review.
If the proposed horizon exceeds the documented source limit, escalating the setting is unavailable within this scope. A separately retained archive is useful only if the chosen recovery mechanism can consume it safely with a defensible boundary. An arbitrary exported log file is not automatically a supported DMS source. Do not solve a storage alert by deleting the only recovery evidence, or solve a CDC alert by allowing the primary database to exhaust storage.
6. Use metrics to locate the delay without inventing continuity
AWS defines CDCLatencySource around capture delay and CDCLatencyTarget around the oldest target work awaiting confirmation. Target latency includes source delay, so adding the two double-counts that component. Their difference is not, by itself, a measurement of the oldest required replay position or of catch-up capacity. The monitoring documentation also describes idle-source resets to zero and transaction-related effects on source latency. A zero point is not a proof of retained history or exhaustive target agreement. Preserve the task dimensions, sample interval, missing samples and activity context with each metric. DMS task metrics
Use a time series to decide where to investigate. Source delay with capture work accumulating calls for source, capture and connecting-path evidence. A relatively current capture side with growing target delay points toward application of changes and its dependencies. AWS's latency guidance recommends investigating endpoints, replication resources and network conditions. It does not supply a universal sustainable rate for the formula above. Troubleshooting DMS latency
Retain both successful and failed measurements. A throughput value collected during an empty target or a quiet source can overstate spare capacity during recovery. Long transactions, full-load competition, validation reads and a different change mix may change the result. These are reasons to request representative authorized observations, not to multiply an untested average by an optimistic factor. If useful capacity cannot be bounded, leave the planned interruption unresolved.
Re-evaluate while recovering. An approach to the oldest available range, an unexpected source replacement, denied log access, storage pressure or stalled capture invalidates the prior plan. Preserve the last known task and source evidence, alert the decision owner and follow the separately approved containment procedure. This article does not instruct an operator to purge files, alter parameters, reload a target or transfer write authority.
7. Keep an owned budget record instead of a green dashboard
The migration owner assembles the following filled synthetic record. It preserves why the base time result cannot approve the full plan. Controlled references should point to real evidence without embedding credentials, raw row payloads or private customer data in the review ticket.
- Record, owner and evidence time
- B017-EXAMPLE-1; fictional migration owner Mira; synthetic time T0, no live observation
- Source, task and endpoint revisions
- SRC-17 RDS for MySQL 8.0 instance; TASK-17 task-based DMS; actual deployed patches and endpoint revisions UNKNOWN
- Recovery boundary and mapping evidence
- Fictional position P17, age 2 hours; source identity and continuous range stipulated; actual provider mapping still required
- Configured and observed log availability
- Proposed 24 hours; fictional required range available; actual configuration and readable inventory UNKNOWN
- Work units and existing backlog
- 24 GiB source-binlog work, independently stipulated; one GiB = 1,073,741,824 bytes
- Arrival and recovery capacity evidence
- r 12 and c 30 GiB/hour, constant synthetic rates; no measured AWS capacity
- Interruption and reserve
- D 6 hours through useful capture; M 4 hours for separate uncertainty; no duplicated allowance
- Capture catch-up and reconciliation
- 96 GiB at restart; 18 GiB/hour spare; 5 hours 20 minutes; 160 GiB processed = 96 backlog + 64 arrivals
- Retention model and time result
- Oldest position held fixed through capture catch-up; 17 hours 20 minutes required; 24-hour policy leaves 6 hours 40 minutes
- Storage basis and result
- Synthetic physical-byte equivalence; 288 GiB retained versus 200 GiB allowance; short by 88 GiB
- Target application and reconciliation
- Not assessed; no target observation, business agreement or cutover acceptance
- Observation access and change authority
- No access exercised; actual approved read identity and separate retention/restart authorizers required
- Combined disposition and invalidation
- HOLD combined proposal; source/range/configuration unknown in reality and synthetic storage fails; changes to any input require re-review
- Owned next evidence and permitted action
- Mira requests authorized source inventory, recovery mapping and representative capacity/storage evidence; no operational action authorized
For a blank record, use the same fields below. Enter UNKNOWN explicitly where the observation is missing. The evidence reference, observed value, responsible owner and observation time belong together; a blank field must not become zero or “available.”
- Record, owner and evidence time
- Enter the review identity, accountable migration owner and time basis.
- Source, task and endpoint revisions
- Enter exact resource identities, deployment type, versions and configuration references.
- Recovery boundary and mapping evidence
- Enter the actual supported recovery position, its age and source-range relationship.
- Configured and observed log availability
- Enter policy readback, required range evidence, gaps and access failures.
- Work units and existing backlog
- Enter quantity, comparable unit, measurement method and covered interval.
- Arrival and recovery capacity evidence
- Enter representative rate bounds, change mix, conditions and observation references.
- Interruption and reserve
- Enter every delay through useful recovery and a separate explained reserve.
- Capture catch-up and reconciliation
- Enter restart backlog, spare capacity, duration and total-work reconciliation, or UNKNOWN.
- Retention model and time result
- Declare fixed or evidenced moving frontier, required horizon and remaining margin.
- Storage basis and result
- Enter physical log-generation evidence, proposed retained volume and approved allowance.
- Target application and reconciliation
- Record independent target progress and acceptance evidence, not an inferred source result.
- Observation access and change authority
- Name scoped read permissions and separate decision owners for changes and recovery.
- Combined disposition and invalidation
- Record time, range, support and storage results separately, plus invalidating changes.
- Owned next evidence and permitted action
- Name the missing artifact, responsible person, review deadline and authorized next step.
Read access can still expose sensitive data or add source load. Ask the database and security owners to approve the minimum observation method, its limits and artifact handling. The documented CDC privileges are not a reason to give every reviewer the migration account. Configuration observation, log inspection, metrics and task readback can have different permissions. A denied observation remains unknown; it must not trigger a broader grant or a production modification by default.
8. Decide what to change before scheduling the interruption
When the range exists and both time and storage screens fit, the next useful action is a separately authorized recovery rehearsal under the actual engine, task and target configuration. The migration owner should agree expected updates, deletes, interruption boundaries and failure conditions before execution. Retain actual source-range and checkpoint observations through that exercise, then reconcile the target. A passing local calculation alone does not approve a real outage, parameter change, resume or cutover.
If storage fails but time fits, have the database owner review approved capacity, a shorter interruption or a different recovery design before changing retention. If capture capacity cannot exceed arrivals, investigate the bottleneck or a bounded reduction in incoming work under application-owner authority. A quiet period can help only when its changed workload and business impact are acceptable and evidenced. It must not become an assumed ability to stop customer writes.
If the required range is missing, stop the continuity claim and preserve evidence before choosing a controlled new baseline or another supported recovery. If its availability is unknown, the migration owner assigns the exact missing inventory or mapping to an authorized observer. Do not close either case from a healthy task label. After any change, replace the affected observations and retain the failed budget with its resolution.
Use the dual-run reconciliation playbook for the owned comparison and authority-transfer process, and the DMS LOB coverage article when large-value preservation is part of acceptance. For this decision, bring one completed retention record to the migration review. Its first unresolved range, rate or storage assumption should determine the next authorized evidence request.