An EC2 Instance Refresh Succeeded. Why Is the Old Code Still Running?

Diagnose the gap between instance-refresh skip matching and installed application identity. Compare mutable bootstrap, pinned releases and the limits of configuration...

Success can mean that nothing needed replacing

An EC2 Auto Scaling instance refresh can succeed without replacing an instance when skip matching finds no relevant configuration change. That does not establish which application package is running. If user data fetches a mutable release location at startup, changing the package at that location does not change the launch-template version of an already running instance.

The immediate question is not whether to repeat the refresh. It is whether the service compared the configuration you intended, and whether every relevant instance has the accepted application artifact. Keep those questions separate before authorizing another replacement wave.

This diagnostic covers one ordinary EC2 Auto Scaling group using the Rolling strategy, one launch template and one approved instance type. The worked example uses a numbered launch-template version and direct AMI ID, with no concurrent scaling, instance protection, Standby instances or warm pool. It excludes EKS managed node groups, ECS deployments, mixed-instances overrides, root-volume replacement and stateful recovery. Those paths need their own applicability checks. AWS exposes both Rolling and ReplaceRootVolume strategies in the instance-refresh model; this article does not transfer behavior between them.

All counts, release labels and observations below are fictional tabletop inputs. No AWS request, instance replacement, customer release or cost measurement was performed for this article.

What skip matching knows, and what it cannot inspect

With an explicit desired configuration, skip matching compares instances with that configuration. Without one, its reference is the configuration last saved on the group. The distinction matters if someone changed the group before starting the refresh: today's group settings are not necessarily the operator's intended prior release. AWS also documents that skip matching cannot detect code fetched by user data, and that a refresh can immediately succeed when the relevant template or instance-type configuration has not changed. See the skip-matching contract.

An immutable template version freezes its recorded configuration, not every external input consumed at boot. A script whose text says “download the current release” can be identical in versions that launch different application code at different times. The same issue can arise with an unpinned package dependency or a startup configuration fetched from a mutable location. These are application-supply-chain risks, not additional AWS matching fields asserted by this article.

Record the intended application identity independently: a retained release manifest, package or container digest, and a trusted running-build observation tied to an instance ID and collection time. A release name returned by a public health endpoint is weak evidence if it is merely a configured string. The application owner should explain how that observation is derived from the installed artifact and what it does not cover, such as dynamically loaded plugins.

This is not a recommendation to replace every instance forever. When configuration references a retained immutable artifact and running-instance evidence confirms it, avoiding unnecessary replacement can reduce disruption. The problem is treating a configuration match as an inspection of installed code.

Work through three unchanged instances

Consider a fictional reporting API with three InService instances, A, B and C. Each was launched from template lt-report, numeric version 8, direct image AMI-base-1 and the same instance type. Its unchanged user-data script fetched a mutable release location. All three installed release R1 at boot. Later, the release owner changes that location to R2, without modifying template version 8 or the group configuration.

The reviewed refresh specifies that same desired template/version and type, with SkipMatching true. In this deliberately fixed cohort, every instance matches, so the modeled replacement count is zero. The expected configuration outcome is success without refreshing the installed application. A separate, stipulated artifact observation still finds R1 on A, B and C. The desired R2 application acceptance therefore fails despite the successful configuration workflow.

Evidence in the fictional caseValueWhat the reviewer may conclude
Refresh referenceExplicit desired lt-report, version 8The reference was not an inferred new version
Current instance configurationA, B and C match version 8 and typeAll three are configuration-matching instances
PreferenceSkipMatching trueMatching instances need not be replaced
Current mutable release locationR2A new download could obtain R2, not proof of installation
Running artifact observationsA: R1, B: R1, C: R1This fictional cohort has not met the R2 release requirement
Tabletop resultZero replacements; workflow successConfiguration success and application failure coexist

Do not infer this count from a real group's desired capacity alone. Instances may launch, terminate or change membership during observation. A real investigation needs the complete relevant cohort and its time interval. If one instance cannot be identified or inspected, its application state remains UNKNOWN. “Two of three are current” is not a completed three-instance release.

In the fictional three-instance case, EC2 Auto Scaling matches unchanged template version 8 and skips all replacements. A separate application observation finds R1 on every instance although R2 is required. Successful configuration refresh does not establish the accepted running artifact.

Solid arrows follow the modeled configuration decision. Dashed arrows carry configuration status and separately stipulated artifact observations into release acceptance. No application traffic, AWS execution or completed runtime inspection is shown.

Capture the actual request before blaming the matcher

The defaults table currently gives skip matching different defaults: enabled in the console and false in AWS CLI or SDK. Auto rollback is disabled in both. Neither a console screenshot from another change nor a remembered CLI default establishes the submitted preference. Save the reviewed request and available returned preferences for this refresh, including whether the desired configuration was supplied.

Use a scoped existing read identity to inspect group configuration, launch-template versions, instance membership and refresh history. Do not grant release permissions simply to investigate. Record account, Region, group identity, refresh ID, request origin and collection times. Retain redacted request evidence without secrets embedded in user data. The launch-template permission guidance describes separate instance-launch and role-passing checks for change operations; being able to describe the group is not authority to launch its replacements.

DescribeInstanceRefreshes returns refresh history for the previous six weeks and supports pagination. Preserve the specific refresh, its reason, times, desired configuration and preferences, not merely the newest successful row. Missing older history is a retention limitation, not proof that no prior refresh occurred. Its workflow fields are not an application package inventory.

Map current group members to EC2 instance observations and the application owner's artifact evidence. Respect pagination and eventual consistency. A stale membership snapshot can misattribute a build result to a terminated instance. Missing permission, a timeout or an incomplete page leaves a gap, not a passing result. This article supplies no mutating command or permission policy.

Choose an intervention from the release contract

The platform and application owners have three materially different options. Each needs separately approved change authority, interruption limits and representative application checks. None follows automatically from a successful diagnostic read.

OptionUseful whenRemaining trade-off
Keep skip matching and make no replacementThe retained matching instances already satisfy the accepted artifact and behavior contractRequires per-instance evidence; service status alone cannot support this choice
Disable skip matching for a reviewed Rolling refreshA startup fetch must run again on this cohortReplaces matching instances too; a mutable source can change again during the wave
Publish a reviewed new numbered template with retained immutable release inputsThe team needs a reproducible desired release rather than a moving boot targetPackaging and compatibility work remain; immutable inputs still need runtime verification

Disabling skip matching can address the skipped-startup symptom, but it cannot guarantee that every download succeeds or that a moving release location remains consistent. A new numbered template pointing to the same mutable location has the same reproducibility weakness. Where practical, retain an image or package manifest that selects the approved artifact, verify its integrity at startup, and report installed identity through an approved diagnostic path. This is an engineering recommendation, not a claim that Auto Scaling performs those checks for you.

Before replacement, agree how old and new workers coexist, how accepted requests drain, and what happens to background jobs. A stateless process may still have consequential in-flight effects. The broader CI/CD rollout guide owns that compatibility decision; this article diagnoses the AWS matching blind spot rather than replacing the release plan.

Fund any temporary capacity and retained recovery resources using actual Region, instance purchase model, volumes and dependencies. Instance-count percentages are not a cost cap or evidence of available capacity. Use the rightsizing change playbook for load and recovery-capacity acceptance, not a generic warmup value copied into every refresh.

A saved configuration is not a retained application release

Rollback has its own prerequisites and time boundary. AWS documents rollback only while a refresh is still in progress, with an explicit desired configuration, a numbered prior template version and no AMI alias from Systems Manager Parameter Store. The saved prior configuration must also be stable enough to launch. A completed refresh requires a new change decision, not a late rollback of that finished operation. See manual and automatic rollback.

Even an eligible rollback restores configuration through replacement instances; it does not resurrect the terminated processes. If the prior template fetches the current mutable package at startup, returning to that template does not establish that R1 will return. Retain the old application artifact and its compatible external-state contract before claiming a reversible release. Pinning the prior numeric version and AMI alone is insufficient when another boot input remains mutable.

Cancellation does not roll back already replaced instances. By default it waits for transitioning launches and terminations; an option can return cancellation status while those transitions continue. Do not confuse a stop request with restored capacity, stopped spend or recovered application behavior. This diagnostic does not prescribe cancellation or rollback commands.

Stop expansion through the approved change procedure on unknown artifact identity, wrong build, unmet application thresholds, broken visibility, lost recovery capacity or unapproved cost exposure. The recovery owner then decides whether a supported configuration rollback, a newly authorized refresh with retained artifacts, or another recovery procedure can meet the application's obligations. Data and completed external effects require their own reconciliation.

Keep a readable evidence record

This populated record is fictional. Its references are illustrative record names, not actual AWS receipts or measurements. It deliberately ends with an unresolved change decision instead of treating a diagnosis as an executed repair.

FieldFilled exampleBlank record to complete locally
Scope and observation intervalReporting API, fictional group G, 12:00 to 12:10 UTCAccount, Region, group identity, interval and observer
Exact refresh requestRecord Q1: Rolling, lt-report version 8, direct AMI-base-1, explicit desired configuration, SkipMatching trueRefresh ID, retained request, strategy, desired configuration and preferences
Cohort and evidence completenessRecord M1: A, B, C; no concurrent membership changes stipulatedComplete member IDs, pagination, membership times and missing records
Intended artifactRecord R2: retained approved R2 manifestApproved release identity, provenance, compatibility and integrity basis
Running observationsRecords O-A/O-B/O-C: R1 on each, stipulatedInstance-linked artifact observations, times, method and blind spots
DiagnosisMatching skips all three; installed-release requirement failsConfiguration finding separate from PASS, FAIL or UNKNOWN application rows
Recovery evidenceUNKNOWN: current mutable source cannot demonstrate retained R1 startupPrior numeric template, direct image, retained boot inputs, tested launch and recovery constraints
Next decisionHOLD a new wave until owners resolve artifact/recovery evidenceNamed application/platform/financial decisions, limits and separate action approval

Keep the completed record in an approved internal store. Do not submit user data, credential material, customer payloads or internal resource inventories through a public contact form. Recheck the record when template aliases, bootstrap dependencies, membership or application requirements change.

Test interpretations, then resolve the first missing record

An independent reviewer can use these unexecuted fixtures to challenge the diagnosis. They are document-level reasoning cases, not an AWS simulator or deployment test.

CaseStipulated factsExpected interpretation
AThree matching version-8 instances run R1; R2 required; skip matching trueZero modeled replacements; application acceptance FAIL
BSame cohort, skip matching falseInstances are replacement candidates; R2 installation remains unproven
CThree matching instances each have trusted R2 and accepted behavior evidenceNo replacement may be justified; retain the complete evidence
DOnly two current instances have trusted R2 evidenceWhole-cohort acceptance UNKNOWN, not a two-thirds pass
EPrior group template is $LatestDo not claim documented instance-refresh rollback eligibility
FPrior template is numeric/direct AMI but downloads a mutable packageConfiguration eligibility does not prove retained application recovery
GRefresh is finishedA new approved change is distinct from rollback of the finished refresh

The next useful step is to join one exact refresh request to one complete instance cohort and its installed-artifact evidence. If the request is missing, resolve that first. If matching is understood but package identity is unknown, improve the approved application observation rather than rerunning an opaque replacement wave.

Bring that record, the first unresolved row and the release/recovery constraints to Ampity's DevOps and SRE service when the team needs help defining an accepted change. The relevant deliverable is a reproducible release decision with evidence, not a promise that every successful refresh shipped the intended code.

Related services