An EC2 Instance Refresh Succeeded. Why Is the Old Code Still Running?
Diagnose the gap between instance-refresh skip matching and installed application identity. Compare mutable bootstrap, pinned releases and the limits of configuration...
Success can mean that nothing needed replacing
An EC2 Auto Scaling instance refresh can succeed without replacing an instance when skip matching finds no relevant configuration change. That does not establish which application package is running. If user data fetches a mutable release location at startup, changing the package at that location does not change the launch-template version of an already running instance.
The immediate question is not whether to repeat the refresh. It is whether the service compared the configuration you intended, and whether every relevant instance has the accepted application artifact. Keep those questions separate before authorizing another replacement wave.
This diagnostic covers one ordinary EC2 Auto Scaling group using the Rolling strategy, one launch template and one approved instance type. The worked example uses a numbered launch-template version and direct AMI ID, with no concurrent scaling, instance protection, Standby instances or warm pool. It excludes EKS managed node groups, ECS deployments, mixed-instances overrides, root-volume replacement and stateful recovery. Those paths need their own applicability checks. AWS exposes both Rolling and ReplaceRootVolume strategies in the instance-refresh model; this article does not transfer behavior between them.
All counts, release labels and observations below are fictional tabletop inputs. No AWS request, instance replacement, customer release or cost measurement was performed for this article.
What skip matching knows, and what it cannot inspect
With an explicit desired configuration, skip matching compares instances with that configuration. Without one, its reference is the configuration last saved on the group. The distinction matters if someone changed the group before starting the refresh: today's group settings are not necessarily the operator's intended prior release. AWS also documents that skip matching cannot detect code fetched by user data, and that a refresh can immediately succeed when the relevant template or instance-type configuration has not changed. See the skip-matching contract.
An immutable template version freezes its recorded configuration, not every external input consumed at boot. A script whose text says “download the current release” can be identical in versions that launch different application code at different times. The same issue can arise with an unpinned package dependency or a startup configuration fetched from a mutable location. These are application-supply-chain risks, not additional AWS matching fields asserted by this article.
Record the intended application identity independently: a retained release manifest, package or container digest, and a trusted running-build observation tied to an instance ID and collection time. A release name returned by a public health endpoint is weak evidence if it is merely a configured string. The application owner should explain how that observation is derived from the installed artifact and what it does not cover, such as dynamically loaded plugins.
This is not a recommendation to replace every instance forever. When configuration references a retained immutable artifact and running-instance evidence confirms it, avoiding unnecessary replacement can reduce disruption. The problem is treating a configuration match as an inspection of installed code.
Work through three unchanged instances
Consider a fictional reporting API with three InService instances, A, B and C. Each was launched from template lt-report, numeric version 8, direct image AMI-base-1 and the same instance type. Its unchanged user-data script fetched a mutable release location. All three installed release R1 at boot. Later, the release owner changes that location to R2, without modifying template version 8 or the group configuration.
The reviewed refresh specifies that same desired template/version and type, with SkipMatching true. In this deliberately fixed cohort, every instance matches, so the modeled replacement count is zero. The expected configuration outcome is success without refreshing the installed application. A separate, stipulated artifact observation still finds R1 on A, B and C. The desired R2 application acceptance therefore fails despite the successful configuration workflow.
| Evidence in the fictional case | Value | What the reviewer may conclude |
|---|---|---|
| Refresh reference | Explicit desired lt-report, version 8 | The reference was not an inferred new version |
| Current instance configuration | A, B and C match version 8 and type | All three are configuration-matching instances |
| Preference | SkipMatching true | Matching instances need not be replaced |
| Current mutable release location | R2 | A new download could obtain R2, not proof of installation |
| Running artifact observations | A: R1, B: R1, C: R1 | This fictional cohort has not met the R2 release requirement |
| Tabletop result | Zero replacements; workflow success | Configuration success and application failure coexist |
Do not infer this count from a real group's desired capacity alone. Instances may launch, terminate or change membership during observation. A real investigation needs the complete relevant cohort and its time interval. If one instance cannot be identified or inspected, its application state remains UNKNOWN. “Two of three are current” is not a completed three-instance release.
Solid arrows follow the modeled configuration decision. Dashed arrows carry configuration status and separately stipulated artifact observations into release acceptance. No application traffic, AWS execution or completed runtime inspection is shown.
Capture the actual request before blaming the matcher
The defaults table currently gives skip matching different defaults: enabled in the console and false in AWS CLI or SDK. Auto rollback is disabled in both. Neither a console screenshot from another change nor a remembered CLI default establishes the submitted preference. Save the reviewed request and available returned preferences for this refresh, including whether the desired configuration was supplied.
Use a scoped existing read identity to inspect group configuration, launch-template versions, instance membership and refresh history. Do not grant release permissions simply to investigate. Record account, Region, group identity, refresh ID, request origin and collection times. Retain redacted request evidence without secrets embedded in user data. The launch-template permission guidance describes separate instance-launch and role-passing checks for change operations; being able to describe the group is not authority to launch its replacements.
DescribeInstanceRefreshes returns refresh history for the previous six weeks and supports pagination. Preserve the specific refresh, its reason, times, desired configuration and preferences, not merely the newest successful row. Missing older history is a retention limitation, not proof that no prior refresh occurred. Its workflow fields are not an application package inventory.
Map current group members to EC2 instance observations and the application owner's artifact evidence. Respect pagination and eventual consistency. A stale membership snapshot can misattribute a build result to a terminated instance. Missing permission, a timeout or an incomplete page leaves a gap, not a passing result. This article supplies no mutating command or permission policy.
Choose an intervention from the release contract
The platform and application owners have three materially different options. Each needs separately approved change authority, interruption limits and representative application checks. None follows automatically from a successful diagnostic read.
| Option | Useful when | Remaining trade-off |
|---|---|---|
| Keep skip matching and make no replacement | The retained matching instances already satisfy the accepted artifact and behavior contract | Requires per-instance evidence; service status alone cannot support this choice |
| Disable skip matching for a reviewed Rolling refresh | A startup fetch must run again on this cohort | Replaces matching instances too; a mutable source can change again during the wave |
| Publish a reviewed new numbered template with retained immutable release inputs | The team needs a reproducible desired release rather than a moving boot target | Packaging and compatibility work remain; immutable inputs still need runtime verification |
Disabling skip matching can address the skipped-startup symptom, but it cannot guarantee that every download succeeds or that a moving release location remains consistent. A new numbered template pointing to the same mutable location has the same reproducibility weakness. Where practical, retain an image or package manifest that selects the approved artifact, verify its integrity at startup, and report installed identity through an approved diagnostic path. This is an engineering recommendation, not a claim that Auto Scaling performs those checks for you.
Before replacement, agree how old and new workers coexist, how accepted requests drain, and what happens to background jobs. A stateless process may still have consequential in-flight effects. The broader CI/CD rollout guide owns that compatibility decision; this article diagnoses the AWS matching blind spot rather than replacing the release plan.
Fund any temporary capacity and retained recovery resources using actual Region, instance purchase model, volumes and dependencies. Instance-count percentages are not a cost cap or evidence of available capacity. Use the rightsizing change playbook for load and recovery-capacity acceptance, not a generic warmup value copied into every refresh.
A saved configuration is not a retained application release
Rollback has its own prerequisites and time boundary. AWS documents rollback only while a refresh is still in progress, with an explicit desired configuration, a numbered prior template version and no AMI alias from Systems Manager Parameter Store. The saved prior configuration must also be stable enough to launch. A completed refresh requires a new change decision, not a late rollback of that finished operation. See manual and automatic rollback.
Even an eligible rollback restores configuration through replacement instances; it does not resurrect the terminated processes. If the prior template fetches the current mutable package at startup, returning to that template does not establish that R1 will return. Retain the old application artifact and its compatible external-state contract before claiming a reversible release. Pinning the prior numeric version and AMI alone is insufficient when another boot input remains mutable.
Cancellation does not roll back already replaced instances. By default it waits for transitioning launches and terminations; an option can return cancellation status while those transitions continue. Do not confuse a stop request with restored capacity, stopped spend or recovered application behavior. This diagnostic does not prescribe cancellation or rollback commands.
Stop expansion through the approved change procedure on unknown artifact identity, wrong build, unmet application thresholds, broken visibility, lost recovery capacity or unapproved cost exposure. The recovery owner then decides whether a supported configuration rollback, a newly authorized refresh with retained artifacts, or another recovery procedure can meet the application's obligations. Data and completed external effects require their own reconciliation.
Keep a readable evidence record
This populated record is fictional. Its references are illustrative record names, not actual AWS receipts or measurements. It deliberately ends with an unresolved change decision instead of treating a diagnosis as an executed repair.
| Field | Filled example | Blank record to complete locally |
|---|---|---|
| Scope and observation interval | Reporting API, fictional group G, 12:00 to 12:10 UTC | Account, Region, group identity, interval and observer |
| Exact refresh request | Record Q1: Rolling, lt-report version 8, direct AMI-base-1, explicit desired configuration, SkipMatching true | Refresh ID, retained request, strategy, desired configuration and preferences |
| Cohort and evidence completeness | Record M1: A, B, C; no concurrent membership changes stipulated | Complete member IDs, pagination, membership times and missing records |
| Intended artifact | Record R2: retained approved R2 manifest | Approved release identity, provenance, compatibility and integrity basis |
| Running observations | Records O-A/O-B/O-C: R1 on each, stipulated | Instance-linked artifact observations, times, method and blind spots |
| Diagnosis | Matching skips all three; installed-release requirement fails | Configuration finding separate from PASS, FAIL or UNKNOWN application rows |
| Recovery evidence | UNKNOWN: current mutable source cannot demonstrate retained R1 startup | Prior numeric template, direct image, retained boot inputs, tested launch and recovery constraints |
| Next decision | HOLD a new wave until owners resolve artifact/recovery evidence | Named application/platform/financial decisions, limits and separate action approval |
Keep the completed record in an approved internal store. Do not submit user data, credential material, customer payloads or internal resource inventories through a public contact form. Recheck the record when template aliases, bootstrap dependencies, membership or application requirements change.
Test interpretations, then resolve the first missing record
An independent reviewer can use these unexecuted fixtures to challenge the diagnosis. They are document-level reasoning cases, not an AWS simulator or deployment test.
| Case | Stipulated facts | Expected interpretation |
|---|---|---|
| A | Three matching version-8 instances run R1; R2 required; skip matching true | Zero modeled replacements; application acceptance FAIL |
| B | Same cohort, skip matching false | Instances are replacement candidates; R2 installation remains unproven |
| C | Three matching instances each have trusted R2 and accepted behavior evidence | No replacement may be justified; retain the complete evidence |
| D | Only two current instances have trusted R2 evidence | Whole-cohort acceptance UNKNOWN, not a two-thirds pass |
| E | Prior group template is $Latest | Do not claim documented instance-refresh rollback eligibility |
| F | Prior template is numeric/direct AMI but downloads a mutable package | Configuration eligibility does not prove retained application recovery |
| G | Refresh is finished | A new approved change is distinct from rollback of the finished refresh |
The next useful step is to join one exact refresh request to one complete instance cohort and its installed-artifact evidence. If the request is missing, resolve that first. If matching is understood but package identity is unknown, improve the approved application observation rather than rerunning an opaque replacement wave.
Bring that record, the first unresolved row and the release/recovery constraints to Ampity's DevOps and SRE service when the team needs help defining an accepted change. The relevant deliverable is a reproducible release decision with evidence, not a promise that every successful refresh shipped the intended code.