EKS AL2023 Replacement: Evidence Before Draining AL2

Build a no-user-launch-template AL2023 replacement admission packet with actual node identity, separate IAM checks, stateless application evidence and a funded...

1. Decide whether the replacement can carry the workload

A replacement node group being active is not permission to drain the old one. The platform operator must connect the requested image family to actual nodes, and the application owner must show that the selected workload behaves correctly on them. This playbook produces that admission record before a separately authorized migration.

The reference is narrow: one existing account, Region and EKS cluster; standard x86_64 AL2 managed nodes originally created without a user-supplied launch template; a new AL2023_x86_64_STANDARD managed group also without one; unchanged control-plane version, instance family and architecture. Use one stateless application cohort with an existing verified IRSA identity, no host networking, no persistent volumes or node-local state, and read-only synthetic fixtures for the first test. Stateless does not mean unable to cause external effects. A web handler that triggers email or payments is not an inert test.

Custom AMIs, user launch templates/user data, self-managed nodes, Karpenter, Auto Mode, Windows, Arm, accelerated instances, control-plane upgrades and stateful migrations need separate procedures. If the old group's creation path cannot be established, mark scope UNKNOWN. A generated template in the account is not by itself proof that someone supplied a template.

All commands and examples below are proposed reference material. No AWS account exercise, resource creation, application deployment, spend or successful recovery was performed for this article. Initial completion means a reviewable decision packet. It does not silently authorize production drain, traffic change or deletion.

2. Record the support boundary and current prerequisites

The EKS AL2 transition FAQ sets the EKS-optimized AL2 image support/publishing boundary at November 26, 2025 and the broader AL2 end-of-support date at June 30, 2026. Both are past on this October 8, 2026 source check. Paying for control-plane extended support does not extend AL2 image support. Retaining old nodes temporarily is an explicitly owned residual risk, not a renewed support entitlement.

The AL2023 migration guide requires blue/green replacement for managed groups using the standard template or a template without a custom AMI ID. This playbook selects only the first branch. It requires VPC CNI version 1.16.2 or later. That minimum is necessary, not sufficient: confirm the selected CNI build supports the actual cluster version and configuration.

The same guide introduces nodeadm but exempts managed groups without a user-supplied template from the custom-AMI metadata change. Do not paste a custom NodeConfig, AL2 bootstrap script or manual nodeadm init into this procedure. A need for such customization changes scope.

Operator output: a dated prerequisite sheet with cluster ARN, account/Region, Kubernetes/platform version, exact old-node versions and image IDs, selected replacement release, CNI version and deployment mode, instance architecture, existing application identity and approved subnet/AZ inventory. Recheck before creation. No universal current Kubernetes version, AMI ID or Region availability is supplied here.

Use DescribeCluster for actual cluster version/platform observations and DescribeAddonVersions for managed-add-on version/architecture/platform compatibility, retaining all relevant pages. AMI version information links component releases; recommended-image retrieval provides a Region/version-specific discovery route. A moving recommended image ID is not the pinned releaseVersion request or proof of what launched. Preserve both discovery time and selected release mapping. Do not add an AMI ID through a custom template to pin it within this no-user-template procedure.

Also retain exact-version DescribeClusterVersions versionStatus and support-end dates from the approved Region. Its older status filter is deprecated. Unsupported, contradictory or unknown support evidence holds the fixed-version rehearsal; hand off any control-plane change to a separately admitted procedure instead of selecting a newer version here.

3. Assign authority and preserve a baseline

The change owner names the platform operator, application owner, security reviewer, financial owner, observer and abort authority. One person may cover several roles, but their decisions remain separate. The observer records raw evidence references, UTC time and collection scope. Do not copy tokens, credentials or customer payloads into the packet.

Before requesting replacement capacity, capture the workload artifact digest, namespace/service account, rollout and probe settings, requests/limits, selectors/affinity/topology, PDB generation/status, dependency paths and synthetic-test acceptance conditions. Record which agents and DaemonSets must run on every replacement node. A node can join while a required logging or security agent fails.

The application owner defines baseline correctness, latency/completion, restart and shutdown criteria at an agreed load, including startup and burst behavior. Pick a test that will detect the real dependency the image change threatens. If an agent reads cgroup paths, its compatibility test differs from an API's ordinary health request. Baselines collected on a quiet pool do not establish peak-load acceptance.

Gate: missing artifact identity, undocumented side effects, mixed architectures, unknown template origin or unowned abort authority means HOLD. Resolve the gap before approving capacity creation. The Kubernetes production-operations framework owns the broader change and recovery packet; this procedure supplies the image-specific evidence, not a replacement for it.

4. Review three IAM callers, not one successful request

The node IAM role guidance separates worker-node API/image-pull requirements from VPC CNI permissions, which may use a dedicated role. Record the actual CNI identity arrangement rather than changing it during this exercise. Application AWS access is a third caller, not evidence that either infrastructure role is correct.

For the selected application, IRSA guidance requires the service-account role configuration and compatible SDK credential behavior. Retain the existing association and trust review, application image/SDK version and expected caller identity. A new pool that works because the application falls back to broader node credentials has not passed identity acceptance. Host-network pods and container-boundary exceptions are reasons to leave this selected cohort, not to claim universal credential isolation.

Node, network-agent and application identities require separate permission and caller evidence. A pass for one identity does not transfer to the others.

Permission comparison, not a credential-flow diagram. The selected application keeps IRSA; the packet records the existing CNI arrangement. No role inheritance or tested isolation is implied.

The security reviewer separately checks observer read access, creation authority and eventual retirement authority using the EKS service-authorization reference. Identify CreateNodegroup, selected node-role passing, resource/tag conditions and applicable organizational restrictions. Kubernetes RBAC is independent of AWS permission to describe a group. This is not a deployable minimum-policy recipe.

Output: caller/action/resource/expected-result rows. Test an allowed read of an approved synthetic fixture through the actual application path, plus a read of a specifically approved out-of-scope synthetic fixture expected to be denied. Record action, caller, resource, exact error or result and time, never secret values. Timeout or DNS failure is not the expected authorization denial. Unexpected allowance or node-role fallback is HOLD; do not broaden policies to make the test green.

For the infrastructure rows, the platform operator retains observed node registration/image-pull results and CNI-specific pod-address/network evidence with the responsible identity arrangement and agent observations. Separate failure of address allocation from application DNS, network policy or downstream authorization failure. A healthy agent process alone does not prove its required AWS operations succeeded. Tests must exercise the configured duties without introducing broader privileges or affecting unrelated workloads.

5. Fund the overlap and bound external scaling

Old and replacement pools coexist during testing. Managed-group guidance charges for provisioned resources, including compute, volumes and other infrastructure; “no additional managed-group fee” does not mean a free replacement. Identify what is incremental, what is existing baseline spend and what can continue after the rehearsal.

The financial owner obtains dated Region/purchase-model rates for the approved unchanged instance type and records root-volume storage, logging, data transfer and fixture-dependency costs. Set maximum additional node-hours, storage exposure, a reserve for diagnosis and a named hard decision deadline. A monitoring alert is not an enforced ceiling. If actual growth exceeds the approved plan, the abort authority stops expansion and decides whether safe funded retention or separately authorized retirement is possible.

Illustrative arithmetic, not AWS rates or a measured bill: four replacement nodes for six hours are 24 additional node-hours. A two-hour funded diagnostic extension adds 8, for 32 total. If the approved allowance is 28 node-hours, the extension does not fit: the owner must stop or explicitly approve a new allowance. Multiplying 32 by a hypothetical 0.20 currency units per node-hour gives 6.40 for compute only, not the total cost. Use real rate evidence before any exercise.

Record autoscaler, repair and other automation that can change either pool. Stored scaling configuration may lag actual capacity. The operator must reconcile live nodes and backing-group observations rather than trusting only a desired-size field. Avoid concurrent template, identity, resource-sizing or application changes that would make the result uninterpretable.

6. Prepare an inert request and exact output mapping

The operator prepares, but does not execute from this article, the reviewed creation specification. CreateNodegroup exposes amiType, Kubernetes version, releaseVersion, node role, subnets, instance types, scaling, labels and taints. Record an explicit supported cluster version and approved available AL2023 release instead of relying on changing latest defaults. No launchTemplate is supplied. The source group remains unchanged by this request review.

Request fieldReview evidence required
Account, Region, cluster and new unique group nameIdentity readback and change record, not a copied development profile
AMI type, exact release, current cluster versionAL2023_x86_64_STANDARD, dated availability evidence for the exact version/Region
Instance type/capacity modelSame approved x86_64 family and architecture; no accidental sizing or purchase-model experiment
Node role and subnetsApproved role ARN and network/AZ inventory, independent CNI/application identities
Scaling boundsInitial test size, maximum funded growth and autoscaler/repair ownership
Labels and quarantine taintDistinct rehearsal selectors; admission-controlled test tolerations; required system-agent compatibility
Root volume, access and tagsApproved storage/remote-access policy and charge attribution; no assumed tag propagation
Launch-template fieldDeliberately omitted; no custom image ID/user data/custom metadata overrides

A quarantine NoSchedule taint only excludes pods without a matching toleration. Existing system agents or broadly tolerating workloads can still land there. Review the actual fleet before creation, then inspect every pod on the new nodes. If the team cannot keep an unapproved workload off the pool, do not call it isolated. Make the application replica fixture outside production Service selectors and suppress scheduled external effects; do not assume a separate namespace is sufficient.

Gate/output: the change owner approves the exact request hash, test containment, separate creation/spend authority and observer plan. This document supplies no creation command or production deployment manifest. Any required custom metadata options or user data leave the no-user-template scope.

7. Join request, group and actual node evidence

After a separately authorized creation, the observer captures the request receipt and DescribeNodegroup output: ARN/name, AMI type, release, version, node role, backing resources, health issues, scaling and status. Then map each Kubernetes node's provider ID to the actual instance and image. A group-level family is not a list of actual AMI IDs, and an image ID is not application compatibility evidence.

These illustrative read-only commands were not run. Confirm the profile and Kubernetes context first; replace each placeholder explicitly. Retain required pagination and complete relevant records, not a single favorable row.

aws sts get-caller-identity --profile EXERCISE_READ_PROFILE
aws eks describe-cluster --profile EXERCISE_READ_PROFILE --region APPROVED_REGION --name APPROVED_CLUSTER
aws eks describe-nodegroup --profile EXERCISE_READ_PROFILE --region APPROVED_REGION --cluster-name APPROVED_CLUSTER --nodegroup-name APPROVED_REPLACEMENT_GROUP
aws eks describe-addon --profile EXERCISE_READ_PROFILE --region APPROVED_REGION --cluster-name APPROVED_CLUSTER --addon-name vpc-cni
kubectl --context APPROVED_CONTEXT get nodes -l eks.amazonaws.com/nodegroup=APPROVED_REPLACEMENT_GROUP -o json
kubectl --context APPROVED_CONTEXT -n APPROVED_NAMESPACE get pods -l APPROVED_FIXTURE_SELECTOR -o json
aws ec2 describe-instances --profile EXERCISE_READ_PROFILE --region APPROVED_REGION --instance-ids APPROVED_MAPPED_INSTANCE_ID

The EC2 instance API supplies instance/image, metadata settings and network observations, with eventual-consistency and scope caveats. Unknown or absent mappings stay UNKNOWN. Reconcile every expected new node, not only Ready nodes. Record node OS/kernel/runtime/kubelet, architecture, capacity/allocatable, labels/taints and condition times alongside the EC2 identity.

DescribeAddon applies when the CNI is an EKS-managed add-on. For self-managed CNI, collect the actual deployed image/version and operator configuration instead; an absent managed-add-on response is not proof of absence or compatibility. Use the add-on response fields as one evidence surface, not universal deployment discovery.

8. Test OS and credentials before ordinary readiness

AL2023 changes cgroup behavior and defaults to IMDSv2. For this no-user-template branch, AWS documents hop limit 1; capture actual metadata settings rather than generalizing the custom-AMI/template defaults. Do not increase the hop limit or restore IMDSv1 just to make a failing application work. Such a change requires separate security and scope admission. These are migration-guide boundaries, not proof that every container is isolated.

The application owner and agent owner use Kubernetes cgroup-v2 guidance to check the exact runtime and monitoring/security-agent builds. Code that reads cgroup files or derives heap/concurrency from container limits needs a direct test. Run the retained artifact with the same limits and synthetic workload profile as the baseline; compare startup, resource detection, memory/restarts, dependency calls and result correctness. Record which checks were actually observed, not just supported by a release note.

Observe the application caller through an approved diagnostic that reports identity, not credentials. Confirm the expected IRSA role and permitted fixture access on replacement nodes. Capture expected-denial evidence separately. A successful image pull proves the pull path, not application AWS access. A passed API call from an administrator's laptop proves neither.

Output/gate: runtime/agent compatibility rows, actual workload caller, allowed and denied fixture results, and artifact/node mapping. Missing evidence, unexpected fallback, memory regressions or absent required agents means HOLD. Do not “fix” the test by changing the application artifact without recording a new baseline and admission iteration.

9. Prove placement and useful service before a drain decision

The platform operator lists all pods on replacement nodes and checks test placement against the approved selectors, architecture, taints, topology, subnet/AZ and resource commitments. Required agents must run where intended. The observer records actual node names for each test replica, so results cannot accidentally come from the retained AL2 pool.

The application owner runs the accepted fixture profile on the replacement, including startup/restart, a bounded representative load and graceful termination in the test environment. Check output correctness and dependency behavior as well as probe/Ready signals. If shutdown can leave in-flight effects, use a synthetic recoverable fixture or stop and redesign the test. The initial no-external-write scope is not silently broadened by a load-test tool.

The retained AL2 pool and new AL2023 pool coexist in one unchanged EKS cluster. New-node, identity and application evidence feeds a pre-drain decision. Missing or failed evidence holds migration; a complete packet permits only the next separately authorized step.

Reference coexistence and evidence boundary. Arrows carry review evidence, not production traffic or automatic migration. Retained capacity is a conditional recovery option, not restored image support.

The next migration plan must independently respect Kubernetes disruption semantics and actual replacement capacity. Do not use forced pod deletion, weaken the PDB or reduce group desired size as a substitute for a bounded supported eviction procedure. The PDB/capacity owner explains why a budget cannot create an eligible host. This playbook does not duplicate its drain exercise or execute a production drain.

Gate: identity/platform/application/placement rows must all pass for the exact cohort, and recovery plus cost must remain funded and feasible. The change owner records ADMITTED FOR NEXT APPROVED STEP, HELD or REJECTED, not “migration complete.”

10. Rehearse classification with fictional evidence

These cases are unexecuted tabletop fixtures, not observed AWS or customer outcomes. Their purpose is to expose favorable-looking incomplete packets.

CaseFictional packetExpected disposition
ANew group active; three nodes Ready; application test never scheduled on themHELD: no replacement application evidence
BCNI minimum satisfied; allowed fixture succeeds using the node role rather than expected IRSA roleREJECTED identity acceptance; contain and investigate, no privilege widening
CCorrect new-node mapping; expected caller; denial is a DNS timeout, not an authorization resultHELD: negative permission result UNKNOWN
DTests pass; old pool now lacks enough eligible capacity to accept the cohort againHELD: no viable retained-pool recovery, regardless of test success
ETests pass; request was reviewed for four nodes but an external scaler expanded to eightHELD: new cost/capacity exposure needs observed reconciliation and approval
FExact artifact and node identity, all owned test rows pass, old capacity and funded overlap remain availableAdmit only the next separately approved migration step; no implied retirement

The financial example above can be reused as a seventh case: a requested extension exceeding allowance is a funding decision, not an automatic longer test. Ask each reviewer which concrete record would change a held case into admissible evidence.

11. Copy the admission packet, leaving gaps visible

Fill this locally in an approved evidence store. Use record hashes for retained metadata/test artifacts, not unsupported checksums of entire running nodes. A link without collection scope, identity and time is not an observation. Never replace blank fields with assumed defaults.

P02 replacement admission record
Change ID / observer / collected-at UTC / next recheck:
Account / Region / cluster ARN / Kubernetes + platform version:
Exact versionStatus / support-end dates / dated availability evidence:
Scope: no user-supplied LT old/new confirmed by source record:
Old group ARN / actual node IDs + AMI IDs / observed versions:
New request hash / receipt / group ARN / amiType / releaseVersion:
New node -> provider ID -> instance -> image mapping record:
Same family + x86_64 / capacity model / subnet-AZ inventory:
CNI managed or self-managed / exact build / compatibility evidence:
Node-role ARN / trust and effective permission evidence:
CNI identity arrangement / permission test evidence:
Application service account / IRSA ARN / SDK / actual caller:
Allowed fixture action-resource-result-time / expected-denial evidence:
Artifact digest / test selectors / production selector exclusion:
Every new-node pod / required agents / containment observations:
OS-kernel-runtime / cgroup test / metadata settings observations:
Placement / capacity / startup-load-correctness / shutdown evidence:
Acceptance thresholds and window / observed results / exceptions:
Old pool actual eligible capacity / route-back artifact / constraints:
Scaler-repair-change controls / time limit / abort owner:
Extra node-hours allowance / actual count-time / other cost evidence:
Evidence row states PASS | FAIL | UNKNOWN (attach rows, do not average):
Decision HELD | REJECTED | ADMITTED FOR NEXT APPROVED STEP:
Platform / application / security / financial decisions and dates:
Open limitations / separate migration approval / retirement owner:
Retirement receipt or funded retained resources / billing follow-up:

Keep a row ledger with question / caller or artifact / source / observed result / timestamp / state / owner / next action. PASS means that specified test passed under its stated conditions, not all future workloads. FAIL means observed nonacceptance. UNKNOWN means unavailable, contradictory, stale or unperformed evidence. Any required FAIL or UNKNOWN prevents admission; averaging scores would hide the critical gate.

12. Stop expansion and preserve a feasible recovery path

Before creation, the safe stop is to retain the read-only packet and resolve gaps. After approved test capacity exists but before migration, stop scheduling additional fixtures, contain external effects and preserve the old production arrangement. The operator must distinguish stopping test expansion from terminating nodes: automatic replacement/repair/scaling may still run, and stopping a process does not end resource charges.

For a later approved migration, a retained-pool reversal is possible only if old nodes remain usable, have eligible capacity, accept the retained application artifact, preserve required identities/network access, and have no incompatible external-state changes to reconcile. Record the actual tested route-back procedure and observation window before migration admission. An old group name, saved manifest or zero-sized pool does not guarantee capacity on demand. AL2's ended support remains a risk even if a temporary reversal works.

Stop on an unexpected principal, unapproved production pod on the new pool, wrong image/architecture, lost observability, unmet application thresholds, blocked placement, ambiguous cohort mapping, cost allowance breach or unusable retained capacity. The named abort authority records containment, actual useful-service recovery and unresolved impact. Do not call a reverted selector or restored Ready count recovered service until the application owner accepts the result.

No control-plane downgrade, custom-template fallback, security bypass, manual node bootstrap or stateful reconciliation is provided here. A scope-changing failure goes to a separate decision, not an improvised command.

13. Retire only under a separate verified decision

Admission leaves cleanup open. The retirement owner inventories remaining pods, required agents, temporary selectors/taints, any shared node-role relationships, resource retention and remaining cost. The managed-group deletion guide documents scale-down and forced termination if draining has not completed after five minutes. Deleting a group is therefore not a safe stand-in for application migration. Its role-mapping warning also requires checking dependencies outside the selected cohort.

The owner requests separate retirement authority for the exact group and confirms that the recovery window has closed through accepted service evidence. This article contains no executable delete, forced drain or scale-to-zero command. Record the actual retirement response, instance and volume disposition, remaining temporary resources and later billing observations. An absent group response alone does not prove every related charge ended.

Done when the packet can support a defensible held/rejected/admitted decision for the exact replacement, each observation has scope/time/owner, the next migration action is separately authorized, and residual risks/cost have an explicit disposition. For an exercised migration, additionally require actual useful-service acceptance, recovery evidence and authorized retirement or funded retention. A tabletop-only packet must say it is unexecuted.

Bring the completed packet and unresolved rows to Ampity's reliability review if an independent assessment is needed. The useful next action is to resolve the first missing admission record, not to declare the cluster upgraded because new nodes appeared.

Related services