Choose an AWS Migration Wave by Dependency, Not Server Count
Select a first AWS migration group using transaction boundaries, measured latency, identity dependencies and recoverability, with a worked example and owner worksheet.
Choose the smallest migration group that can complete its business operation and recover safely in the proposed placement. Every dependency left outside that group needs evidence for its latency, data consistency, identity, permissions and failure behavior. An easy-to-copy server is not necessarily a safe migration unit. An unknown critical dependency is a reason to hold the boundary decision, not to label the application independent.
This article helps a service owner and migration lead select a first group before they commit to a wave date. The output is a reviewed dependency-boundary record: what moves together, what stays elsewhere, why each separation is acceptable, and what would invalidate that decision. The example is fictional. Its timings are stipulated teaching inputs, not AWS benchmarks or customer results. No infrastructure was deployed or tested for this article.
Start with an operation, not an inventory row
A server inventory answers where software runs. It rarely explains which components must succeed together when a customer confirms an order, an employee signs in, or a billing job posts a transaction. Start with one such operation and trace it through completion, including background work that makes the visible response truthful.
For example, an API can return “order confirmed” while a worker later reserves stock. If that worker cannot reach its database after the move, the API's availability is not evidence that the business operation works. Conversely, a reporting viewer might consume an immutable nightly export and have no role in accepting orders. Keeping the viewer in the source environment can be acceptable even though it appears on the same application diagram.
AWS distinguishes location-sensitive hard dependencies from dependencies that can operate across a separation. Its portfolio wave-planning guidance also warns that communication data alone does not establish latency tolerance and that no single discovery source captures every relationship. Use network observations to find candidate edges, then ask owners what those edges mean.
Record at least the normal interactive path, scheduled processing, failure handling, and recovery path. An observation window that misses month-end settlement, certificate renewal or a rarely used restore operation cannot certify their absence. Write the coverage gap explicitly and obtain representative evidence before relying on the boundary.
Classify the boundary before choosing the group
“These systems communicate” is too broad to guide a move. Classify the reason a dependency constrains placement. One edge can have several reasons, and the most restrictive unresolved one governs the candidate boundary.
| Boundary question | Evidence needed before separating it |
|---|---|
| Must several reads or writes complete within one business transaction? | Transaction ownership, consistency requirements, commit and retry semantics, and the behavior after a partially observed result |
| Does the caller make serial remote requests? | Representative traces, call count, connection behavior, measured end-to-end timing and an owner-approved latency budget |
| Does successful work depend on an external identity service? | Login and token-refresh paths, key refresh, expiry and revocation behavior, dependency outage behavior and an accepted authorization policy |
| Can the target principal perform the operation and recover it? | Explicit target identity, allowed actions, network path, keys and restore access, plus negative permission checks |
| Can the consumer tolerate delayed or duplicated delivery? | Freshness limit, durable delivery, idempotency, ordering where required, backlog capacity and reconciliation |
| Must components change or recover in one maintenance window? | Named operators, shared change constraints, restoration sequence and available recovery capacity |
Must several reads or writes complete within one business transaction?
Evidence needed before separating it: Transaction ownership, consistency requirements, commit and retry semantics, and the behavior after a partially observed result
Does the caller make serial remote requests?
Evidence needed before separating it: Representative traces, call count, connection behavior, measured end-to-end timing and an owner-approved latency budget
Does successful work depend on an external identity service?
Evidence needed before separating it: Login and token-refresh paths, key refresh, expiry and revocation behavior, dependency outage behavior and an accepted authorization policy
Can the target principal perform the operation and recover it?
Evidence needed before separating it: Explicit target identity, allowed actions, network path, keys and restore access, plus negative permission checks
Can the consumer tolerate delayed or duplicated delivery?
Evidence needed before separating it: Freshness limit, durable delivery, idempotency, ordering where required, backlog capacity and reconciliation
Must components change or recover in one maintenance window?
Evidence needed before separating it: Named operators, shared change constraints, restoration sequence and available recovery capacity
Treat these as evidence questions, not automatic instructions to move everything together. A shared database may serve unrelated workloads with independent transactions. A common identity platform may be a prepared foundation for many waves. AWS's large-migration wave-planning guidance separates move groups from waves and cautions against letting a common shared service join an entire estate into one group. That does not make the shared service optional. Its readiness remains a prerequisite.
For each edge, record one disposition: move together, separate with demonstrated contract, prepare shared foundation first, or hold pending evidence. “Probably fine over the network” is not a fifth disposition.
Work the first-wave example all the way through
Consider a fictional order application with an Intake API, an Order Worker and PostgreSQL in the source environment. A shared identity service supplies authentication. A Reporting Viewer consumes a nightly immutable export. The proposed target is AWS, but this decision does not depend on a specific compute service.
The service owner, Maya, defines the operation as “accept an order and record its reservation status without losing or duplicating an accepted request.” Database owner Leon confirms that the API and worker both participate in the same order state machine. Security owner Priya owns identity and target permissions. Migration lead Omar owns the group proposal and evidence gaps. These names are fictional role examples, not Ampity personnel or customer references.
One representative API path contains eight sequential database round trips. For a preliminary arithmetic screen, assume 120 ms of fixed application and database execution work, excluding those round trips. Source round-trip time is stipulated as 1 ms; a proposed cross-cloud path is stipulated as 25 ms. Maya's illustrative end-to-end budget is 250 ms.
| Arithmetic screen | Result |
|---|---|
| Source placement: 120 + 8 × 1 | 128 ms |
| API-only move: 120 + 8 × 25 | 320 ms |
| Transport allowance: 250 − 120 | 130 ms |
| Maximum round-trip input under this simplified model: 130 ÷ 8 | 16.25 ms |
| API-only modeled excess: 320 − 250 | 70 ms |
Source placement: 120 + 8 × 1
Result: 128 ms
API-only move: 120 + 8 × 25
Result: 320 ms
Transport allowance: 250 − 120
Result: 130 ms
Maximum round-trip input under this simplified model: 130 ÷ 8
Result: 16.25 ms
API-only modeled excess: 320 − 250
Result: 70 ms
The API-only candidate fails this simplified budget screen. Moving the API, worker and database together avoids that particular cross-boundary database path, but does not prove the target meets the budget. The target's query execution, storage, connection setup, encryption, queuing and other calls still need measurement.
Do not call 320 ms a predicted p95. Adding percentile observations does not generally produce an end-to-end percentile. Here, the numbers are deterministic assumptions, with no contention, packet loss or retry cost. Validate the actual journey under representative concurrency, data volume, cold connections and failures. Record its latency distribution and errors against the approved objective. A pass in the arithmetic model is permission to investigate, not acceptance to migrate.
The reporting edge is different. Its owner accepts up to 24 hours of freshness, requires a complete versioned export and can retry delivery without posting orders. It can remain outside the first group if that contract, target export permissions and recovery behavior are demonstrated. If it secretly queries live order tables during business hours, the evidence changes and the classification must be reopened.
Identity also remains conditional. In this fictional design, the API validates some tokens locally using cached keys. A successful request with a warm cache says little about new login, expired tokens, key rotation or loss of the identity endpoint. Priya must establish the accepted behavior for each path. The wave remains on hold if that evidence is unavailable. It does not automatically absorb the entire shared identity estate.
Compare complete candidates, including the smaller alternative
The first candidate need not be the most commercially important application. It should provide useful migration learning within the team's demonstrated capacity, without hiding critical dependencies to make the server count look small.
| Candidate | Boundary finding | Decision at this stage |
|---|---|---|
| Intake API alone | Serial database path fails the fictional timing screen; worker and write-state behavior remain unresolved | Hold; do not schedule this split |
| Intake API, Order Worker and PostgreSQL | Keeps the critical order-state path together; identity, target behavior and recovery still require evidence | Candidate group, not yet migration-ready |
| Reporting Viewer alone | Potentially independent immutable-export contract; lower coverage of the core order workflow | Alternative pilot if its delivery contract is verified and the larger group exceeds capacity |
Intake API alone
Boundary finding: Serial database path fails the fictional timing screen; worker and write-state behavior remain unresolved
Decision at this stage: Hold; do not schedule this split
Intake API, Order Worker and PostgreSQL
Boundary finding: Keeps the critical order-state path together; identity, target behavior and recovery still require evidence
Decision at this stage: Candidate group, not yet migration-ready
Reporting Viewer alone
Boundary finding: Potentially independent immutable-export contract; lower coverage of the core order workflow
Decision at this stage: Alternative pilot if its delivery contract is verified and the larger group exceeds capacity
*Arrows mean database request/reply, not replication or automatic recovery. The group boundary describes migration scope, not an AWS network design. Removing the illustrated split does not establish target performance or migration readiness. Identity and reporting dependencies are omitted from this focused view, not waived; the candidate table and records E-01, E-02 and G-01 retain their evidence requirements.*
If the coherent order group is too large, choose another representative pilot or explicitly fund and test a boundary change. Do not move the API alone simply because the three-component group is inconvenient. Database compatibility, data-copy duration and write-authority transfer remain separate readiness work. Use the database migration guide for that later decision.
Also document what the smaller pilot cannot teach. A reporting viewer may prove routing, deployment and export delivery, but not transactional database cutover. Its success cannot be reused as evidence that an order-writing application is ready.
Know when separation is a better design
Co-location is not the objective. A stable boundary can make migrations smaller and reduce coupled failure. A service with one bounded remote request might meet its objective across a validated path. An asynchronous consumer may tolerate a delivery delay. A reporting workload may use an explicit snapshot rather than a live transactional database. Each case needs its own operating contract.
AWS Reliability Pillar guidance on loose coupling describes asynchronous interaction where an immediate response is unnecessary, together with durable intermediaries and controlled interfaces. Applying that principle requires business agreement about when work is complete. Returning “queued” is not equivalent to “reservation confirmed.”
A queue is not a shortcut around an atomic write requirement. If the application updates a database and publishes an event separately, one can succeed while the other fails. The transactional outbox pattern addresses that dual-write problem by recording the change and outbox entry in one transaction, then relaying it. Duplicate delivery still requires consumer handling. Introducing this pattern is an application change with its own acceptance work, not a network setting in the migration plan.
Similarly, retries cannot repair a boundary that consumes more time than the business allows. They may amplify load or repeat a non-idempotent operation. AWS retry guidance calls for bounded retries, backoff, jitter and explicit testing. Establish which layer retries and what happens when the result of a write is unknown before allowing the split.
Ask for evidence that could reject the proposal
For the order group, the performance check is a traced business journey, not a ping. Capture request identity, relevant spans, serial call count, connection state, concurrency, data shape, duration distribution and errors in both proposed placements. If these differ materially between runs, explain the difference or repeat the comparison. A synthetic idle path cannot close a peak-load question.
The failure check interrupts the permitted cross-boundary dependencies in an authorized rehearsal environment. Observe identity unavailability, expired credentials, export interruption and recovery after a connection break. For a timed-out order write, establish whether the operation committed before deciding to retry. Include denied access, not only successful privileged calls. Never introduce faults into production without its normal change authorization.
The data check follows a stable business identifier through API acceptance, database state and worker outcome. A count of rows or HTTP successes alone cannot distinguish missing, duplicated or wrongly associated work. The recovery check establishes who can restore the necessary state, with which keys and permissions, and how accepted work is reconciled. It is not enough that a backup exists.
Hold the candidate when a critical edge has unknown observation coverage, unowned authorization semantics, an unmet timing objective, ambiguous write outcomes, missing recovery access or an unavailable operational owner. A temporary successful test does not override a known unacceptable failure mode. Reopen the boundary after material changes to request count, identity policy, payload size, topology or business freshness requirements.
Keep one owner worksheet for the decision
Use one record per directed dependency edge, plus a group-level approval record. A dependency graph without these records is a useful map, but not the migration decision. Keep identifiers and sanitized artifact references here, not credentials, customer data or raw sensitive traces.
| Worksheet field | What the owner supplies |
|---|---|
| Decision and edge identity | Group revision, operation, caller, dependency, direction, environment and accountable service owner |
| Business contract | Completion meaning, transaction boundary, freshness allowance, error tolerance and approved objective |
| Placement and observed behavior | Current and proposed location, protocol, serial calls, payload class, connection reuse and representative observation window |
| Identity and permission boundary | Target principal, allowed actions, key access, expiry and denial behavior, with security owner's evidence reference |
| Performance and failure evidence | Traces and test conditions, actual journey results, dependency-loss behavior, retry ownership and unresolved gaps |
| Recovery evidence | Restore and reconciliation path, accepted-work identifier, required artifacts, recovery owner and rehearsal reference |
| Disposition and sign-off | Move together, verified separation, shared-foundation prerequisite or hold; named reviewer, reason, checked date and reopening trigger |
Decision and edge identity
What the owner supplies: Group revision, operation, caller, dependency, direction, environment and accountable service owner
Business contract
What the owner supplies: Completion meaning, transaction boundary, freshness allowance, error tolerance and approved objective
Placement and observed behavior
What the owner supplies: Current and proposed location, protocol, serial calls, payload class, connection reuse and representative observation window
Identity and permission boundary
What the owner supplies: Target principal, allowed actions, key access, expiry and denial behavior, with security owner's evidence reference
Performance and failure evidence
What the owner supplies: Traces and test conditions, actual journey results, dependency-loss behavior, retry ownership and unresolved gaps
Recovery evidence
What the owner supplies: Restore and reconciliation path, accepted-work identifier, required artifacts, recovery owner and rehearsal reference
Disposition and sign-off
What the owner supplies: Move together, verified separation, shared-foundation prerequisite or hold; named reviewer, reason, checked date and reopening trigger
In the fictional record, Maya owns the order contract, Leon owns the database and restoration evidence, Priya owns identity and permissions, and Omar assembles the proposal. A person can hold several roles, but no role should silently inherit another's approval. Technical readiness is not permission to change a contract or delete the source.
Here is a filled example of that record. These are fictional planning entries, not trace captures, test results or signed approvals. UNKNOWN is an unresolved gate, not a zero-risk value. The identifiers only connect the example's records.
| Directed edge record E-01 | Fictional planning entry |
|---|---|
| Caller to dependency | Intake API to Order Database; request and response path for confirming an order |
| Proposed split and owner | API moves to AWS; database stays at source. Maya owns the operation; Leon owns database behavior. |
| Business and timing contract | Complete the order-state operation within the illustrative 250 ms budget; eight sequential database round trips; fixed work excludes transport. |
| Evidence status | Stipulated screen only: 120 + 8 × 25 = 320 ms, exceeding the budget by 70 ms. Representative end-to-end measurements, transaction outcomes, target permissions and recovery evidence are UNKNOWN. |
| Disposition | Reject the API-only split under these assumptions. Carry API, worker and database forward as a candidate group, not an accepted migration. |
| Reopening trigger | Change in serial call count, timing budget, measured path, database transaction design or placement. Recompute the screen and obtain the missing evidence before reversing the decision. |
Caller to dependency
Fictional planning entry: Intake API to Order Database; request and response path for confirming an order
Proposed split and owner
Fictional planning entry: API moves to AWS; database stays at source. Maya owns the operation; Leon owns database behavior.
Business and timing contract
Fictional planning entry: Complete the order-state operation within the illustrative 250 ms budget; eight sequential database round trips; fixed work excludes transport.
Evidence status
Fictional planning entry: Stipulated screen only: 120 + 8 × 25 = 320 ms, exceeding the budget by 70 ms. Representative end-to-end measurements, transaction outcomes, target permissions and recovery evidence are UNKNOWN.
Disposition
Fictional planning entry: Reject the API-only split under these assumptions. Carry API, worker and database forward as a candidate group, not an accepted migration.
Reopening trigger
Fictional planning entry: Change in serial call count, timing budget, measured path, database transaction design or placement. Recompute the screen and obtain the missing evidence before reversing the decision.
| Directed edge record E-02 | Fictional planning entry |
|---|---|
| Caller to dependency | Intake API to shared identity endpoint for the applicable login, token or key-refresh paths; local validation is a separate assumed path |
| Proposed split and owner | Order group moves; shared identity stays. Priya owns authorization semantics and permitted target access. |
| Required contract | Establish new-login, token expiry, key refresh, revocation and endpoint-outage behavior against the service owner's accepted policy. A warm cached-token request is insufficient. |
| Evidence status | All required identity and recovery checks are UNKNOWN. No successful rehearsal or approved degraded mode is recorded. |
| Disposition | HOLD order-group readiness. Treat identity as a shared-foundation prerequisite; do not absorb the whole identity estate merely to close the diagram. |
| Reopening trigger | Identity policy, token lifetime, refresh implementation, target principal, permitted network path or recovery access changes. |
Caller to dependency
Fictional planning entry: Intake API to shared identity endpoint for the applicable login, token or key-refresh paths; local validation is a separate assumed path
Proposed split and owner
Fictional planning entry: Order group moves; shared identity stays. Priya owns authorization semantics and permitted target access.
Required contract
Fictional planning entry: Establish new-login, token expiry, key refresh, revocation and endpoint-outage behavior against the service owner's accepted policy. A warm cached-token request is insufficient.
Evidence status
Fictional planning entry: All required identity and recovery checks are UNKNOWN. No successful rehearsal or approved degraded mode is recorded.
Disposition
Fictional planning entry: HOLD order-group readiness. Treat identity as a shared-foundation prerequisite; do not absorb the whole identity estate merely to close the diagram.
Reopening trigger
Fictional planning entry: Identity policy, token lifetime, refresh implementation, target principal, permitted network path or recovery access changes.
Group record G-01, assembled by Omar: the candidate contains Intake API, Order Worker and Order Database. E-01 rejects the smaller API-only split; it does not accept the larger group. E-02 remains HOLD. Target performance, order-state reconciliation, worker behavior, restoration access and delivery capacity are still UNKNOWN. The Reporting Viewer remains a separate potential pilot only after its immutable-export contract is verified. Maya has not signed readiness, and no wave date is committed. Omar's next action is to obtain the named owners' missing evidence and issue a new record revision, not convert UNKNOWN fields into approval.
Before scheduling, the service owner confirms that every critical external edge has a reviewed disposition, the proposed group fits delivery and recovery capacity, and the hold conditions are closed with evidence. Record unresolved noncritical limitations rather than hiding them. Only then feed the group into wave scheduling. Renewal deadlines, database cutover and source retirement have their own gates; selecting a coherent group does not satisfy them.
Attach the signed group record to the wave proposal. Keep the rejected API-only option and its 70 ms modeled excess in that record so a later scope reduction cannot silently recreate the same unsafe split.