ECS Rollback Needs an Eligible Completed Deployment

Diagnose missing automatic rollback targets in rolling ECS deployments with a dated completed-state worksheet, temporary eligibility boundaries and explicit...

Enabling ECS deployment circuit-breaker rollback does not create a previous successful deployment. For the rolling deployment scope described here, automatic rollback needs an eligible deployment in COMPLETED state. A first failed deployment can have no such target. A revision that was completed can also become temporarily ineligible while its rollback is in progress.

Check the exact service configuration and dated deployment states before treating the rollback flag as a recovery plan. This article gives ECS operators a worksheet to distinguish an evidenced target, an absent target and an unknown target. Examples are fictional, not captured AWS responses. No account was queried, deployment attempted, event rule enabled or failure injected. This is a diagnosis, not a production mutation procedure or proof that application data can be rolled back.

1. Establish which deployment mechanism you actually have

Record the account, Region, cluster and service, then the deployment controller and strategy. Do not identify the mechanism from a Fargate label or from the fact that a service runs containers. Capture the circuit breaker's enable and rollback values separately. The DeploymentCircuitBreaker API distinguishes those fields and specifically limits its applicability to rolling ECS deployments.

There is a current documentation discrepancy. The broader failure-detection page lists rolling and blue/green support for both methods, while the specific API and detailed breaker guide restrict this mechanism to rolling ECS. This article does not resolve that conflict through an invented account test. Its worked scope is rolling deployment using the ECS controller, with applicable state evidence. An ECS controller alone is not enough to assume that every strategy has the same behavior. For other configurations, hold this worksheet's applicability until the exact mechanism is established.

Retain the current configuration reference rather than a template that may not have reached the service. If rollback is disabled, target eligibility alone cannot establish automatic rollback. If the setting is unreadable, record UNKNOWN, not enabled by convention. Do not change a service or widen permissions simply to make the diagnostic record complete.

2. Separate deployment status from completion evidence

A service's active status, a deployment's PRIMARY designation and running tasks answer different questions from completed rollout. The Deployment schema defines PRIMARY/ACTIVE/INACTIVE separately from rolloutState. It also limits returned rollout state to rolling ECS services not behind a Classic Load Balancer. An absent field outside that scope is not a failed or completed value.

Record each relevant deployment ID, task-definition reference, created/updated time, state and reason with the collection time. Include the service and configuration reference from the same context. A retained screenshot of yesterday's completed state does not establish today's eligibility after another operation began. Keep a historic state as history, not the current answer.

The DescribeServices reference exposes service configuration, deployments and failures in the response. Use explicit cluster/service context in an authorized readback. An HTTP success or a returned service name does not erase per-item failures or establish a complete history. Missing access, an incorrect context or unavailable relevant state leaves an evidence gap. Do not manufacture an empty candidate set from that gap.

3. Follow the candidate's state, not only the failed release

The detailed circuit-breaker guide says rollback selects the most recent completed deployment. When rollback starts, that selected deployment changes to IN_PROGRESS and is not eligible for another rollback until it completes again. When no completed deployment is found, the breaker does not launch new tasks and the deployment stalls.

The consequence is a temporary eligibility boundary. “D1 completed earlier” and “D1 is completed now” can both be relevant facts with different answers. If D1 is the only completed candidate and is now in progress, an additional deployment cannot rely on its earlier state as an automatic escape path. If another completed candidate exists, it must be assessed from the actual relevant history rather than assumed away. The diagram deliberately stipulates that no other completed revision exists.

Three fictional snapshots follow D1 from completed and eligible, through rollback in progress and temporarily ineligible, to completed again. The middle snapshot has no other completed revision. Arrows mean state transitions, not task traffic or recovered data.

*Conditional reference: D2 has failed, rollback is enabled, and D1 is the only candidate. Blue arrows show D1 state changes; they do not promise elapsed time, successful rollback or business recovery. The final completed state is a stipulated scenario, not an observed result.*

Do not queue another speculative release to discover whether the target remains usable. First preserve the failed release and selected candidate identities, then establish their current states. A new operation can change the very evidence the operator is trying to interpret.

4. Complete a hypothetical eligibility worksheet

The rows below stipulate applicable rolling ECS configuration, breaker enabled and rollback enabled, except where stated. Deployment labels and snapshot order are fabricated. They are not account output, execution results or an exhaustive ECS state machine. “Eligible target” means the narrow documented prerequisite is represented; it does not mean recovery succeeded.

PacketRelevant snapshotNarrow worksheet dispositionNext owner/action
A: first deploymentD1 FAILED; no earlier COMPLETED deployment in the stipulated complete setNo eligible automatic target evidencedService owner diagnoses the failed launch and chooses an authorized recovery intervention
B: later failureD1 COMPLETED; newer D2 FAILEDD1 is the eligible candidate in this exampleOperator verifies actual selection/start and separately checks application recovery
C: rollback startedD1 IN_PROGRESS after selection; D2 FAILED; no other COMPLETEDD1 temporarily ineligible; no eligible target in this snapshotRelease owner holds the next progression and observes the current recovery operation
D: recovery revision completesD1 COMPLETED again; D2 FAILEDD1 is eligible again under the stipulated evidenceOperator retains the new completion evidence, not the old screenshot
E: missing stateD1 PRIMARY; rolloutState unavailable or deniedUNKNOWN, not completed and not proven absentEvidence owner resolves applicability/access/current state
F: rollback disabledD1 COMPLETED; D2 FAILED; rollback=falseCandidate state does not establish automatic rollbackConfiguration owner establishes the intended mechanism without silently enabling it

Packet B is not permission to select a preferred revision manually while calling it the automatic target. Packet C's absence conclusion depends on its stated complete set; an incomplete export would instead be UNKNOWN. Packet D is not a claim that rollback always completes. Packet E prevents a familiar operational shortcut: interpreting a status label or lack of evidence as a successful baseline.

The local validator checks these stipulated distinctions, including choosing the latest completed candidate rather than the latest deployment of any state. It does not reproduce ECS failure detection, task scheduling, timestamps or the service's historic retention. No locally passing fixture becomes an AWS result.

5. Retain a reusable evidence record

Copy this record for one service and one decision time. Broadly shared versions should use aliases; preserve exact identifiers in the permitted operational source. Leave unresolved fields UNKNOWN and assign an evidence owner. Do not substitute the fictional D1/D2 states above.

Record ID / decision time / service owner / evidence collector:
Account / Region / explicit cluster / service / collection time:
Deployment controller / strategy / load-balancer applicability:
Current configuration reference / breaker enable / rollback:
Relevant deployment IDs / task-definition refs / created and updated times:
Each current rolloutState / reason / separate PRIMARY or ACTIVE status:
Complete relevant-set evidence / omitted or unavailable entries:
Most recent COMPLETED candidate / supporting current record or UNKNOWN:
Selected rollback deployment / actual start reference / current state:
Failed deployment / reason / retained event or service evidence:
Unresolved source discrepancy / denied read / missing field / owner:
Application/data compatibility evidence / separate recovery authority:
Disposition: candidate evidenced / none evidenced / UNKNOWN / out of scope:
Next action / accountable owner / stop condition / next readback time:
Actual AWS result: NOT EXECUTED unless observed and referenced:

The useful output is an explainable disposition at a recorded time. Keep the source result and the analyst's interpretation separate. If a later readback changes the candidate, revise the decision record and retain the earlier version. Do not make an apparently stable worksheet by overwriting every historic state with the newest one.

6. Corroborate rollback without declaring the application recovered

The deployment-event reference distinguishes in-progress, completed and breaker-failed deployment events; in-progress includes initial and rollback deployments. Retained events can corroborate timing and identity when available. This article neither configures an event rule nor assumes a complete event archive. An event you did not collect is not evidence that the operation did not happen.

Use deployment identities, reason and dated service readback together rather than a notification headline alone. Investigate mismatches before declaring that the expected revision was selected. A completion event is a scheduler-related fact, not proof that old code can read newly written data or that pending business work is reconciled. Keep the compatibility and effect evidence owned by the feature-flag and data-state rollback article.

The cloud deployment acceptance pack owns the broader release record, authority and acceptance checks. Add this worksheet as its ECS-specific eligibility evidence, not as a replacement. A record with an eligible target but unknown application compatibility can remain held for recovery use. Conversely, compatible old code cannot invent a missing completed scheduler target.

7. Make the missing-target response an owned decision

If there is no evidenced candidate, the result is not “keep retrying until rollback works.” The service owner needs to identify the current failure and choose a bounded recovery intervention with the release authority. Depending on the actual fault and current state, that decision might concern correcting a launch prerequisite or preparing a separately validated forward revision. These are possible operator decisions, not executable instructions or a promise that either resolves this incident.

Do not disable the breaker, alter desired count, recreate the service or initiate another deployment merely to clear a status without understanding the resulting state. Such actions can affect capacity, identity, exposure and evidence. This diagnostic article authorizes none of them. An existing completed revision is also not a universal reason to force rollback when current data or external effects are incompatible.

Before a separately authorized rehearsal, define the service scope, synthetic inputs, permitted actions, observer, budget, abort conditions and cleanup evidence. Exercise the expected no-baseline case and the temporarily-in-progress case without customer traffic. Keep the expected answer separate from actual results. This educational plan does not report an executed AWS rehearsal. Before implementation, the workload owner must review the scoped recovery decision.

8. Check one service before relying on automatic recovery

Start with the current controller/strategy and separate breaker settings. Establish whether rollout state is applicable and readable, retain the relevant dated deployment set, and identify the most recent completed candidate or mark the conclusion UNKNOWN. Recheck the selected revision when rollback begins rather than treating its earlier completion as permanent eligibility.

That packet answers a narrow question: does this service have an evidenced automatic rollback target under the supported mechanism? It does not answer every release-health, data-recovery or business-acceptance question. Take the resulting gaps to the named release and evidence owners before the next progression. No demand forecast, uptime promise or deployment success is implied by a completed worksheet.

Related services