Replace Static CI Cloud Credentials Without Hiding a Fallback
Migrate one deployment identity with observed allow-and-deny tests, a controlled cutover and separate evidence for retiring keys, sessions and stored copies.
trigger="A deployment job uses a stored cloud key, and the team needs a replacement whose identity, permissions and retirement evidence can be inspected." owner="The platform owner accountable for deployment continuity and the completed migration record." participants={['Workflow maintainer', 'Cloud identity administrator', 'Security reviewer', 'Service owner', 'Release approver']} prerequisites={['Approved account and consumer inventory', 'Isolated role and harmless test target', 'Independent recovery access and stop owner', 'Restricted evidence storage without tokens or secret values']} outputs={['A consumer and credential-source register', 'A versioned identity and permission contract', 'Observed positive and negative test results', 'A cutover and recovery record', 'Separate key, session and copy disposition evidence']} doneWhen={['Required jobs use the intended identity', 'Forbidden contexts fail at the right layer', 'Legacy access and sessions are resolved', 'Stored copies have an approved disposition', 'Recovery access is tested and owned']} />
Migrate one access path before declaring the key retired
Use this procedure to replace a long-lived CI cloud credential with a workload identity while preserving deployment continuity. Finish with an evidence packet that identifies which jobs migrated, which identities are accepted, which resource actions are permitted and what happened to the old access path. A green deployment is only one result in that packet. It does not establish that the old key stopped working or that an unintended job cannot obtain the new role.
The running example is a hypothetical repository with beta and production release jobs. It uses GitHub Actions and AWS to make the review boundaries concrete. No Ampity hosting configuration, account access or customer migration is claimed here. Other CI and cloud providers require their own documented claims, credential precedence and withdrawal behavior. This playbook is a proposed operating procedure, not authorization to change a production identity.
Read the workload identity article for the trust-policy boundary, and the credential retirement article for the closure evidence. This procedure connects both into a controlled migration. If the credential may already be compromised, use the incident process and its containment authority instead of waiting for a routine cutover window.
These are proposed change-acceptance gates, not a token exchange topology. Complete the consumer inventory and recovery planning before entering them. Missing evidence holds progression; the diagram does not authorize reactivating a retired key or reversing an already completed deployment.
1. Define the migration boundary and independent recovery access
Owner: platform owner. Output: approved migration contract. Name the credential identifier, cloud account, repositories, environments and required consumer jobs. Record who can approve the identity change, who can stop a run and who accepts the service result. Preserve the credential's identifier in restricted evidence, not its secret value in a ticket. A repository name alone is insufficient when several jobs share an organization secret or runner configuration.
Define the maximum disruption and the conditions that halt cutover. Examples include an unexpected principal, a forbidden action succeeding, an unowned dependent job, loss of recovery access or evidence exposing a token. The approver must choose the actual tolerances for the service. Do not inherit a universal acceptable outage window from this example or treat normal deployment access as permission to administer identity policy.
Establish an independent, authorized recovery path and test it against a harmless target. If the same CI role is both the object of the change and the only means of repairing it, a denied exchange can lock out the team. Name the recovery owner and record the required control-plane access. Recovery access should be constrained and auditable; it is not a shared administrator key placed beside the credential being retired.
2. Inventory consumers and every credential source they can select
Owner: workflow maintainer. Output: consumer and source register. Include normal release jobs, scheduled integration jobs, manual operations, reusable workflows and disaster-recovery scripts. Record triggers, owner, last observed execution, required operations and the credential source used by each. Mark unknown consumers explicitly. An inventory derived only from recently successful builds can miss a monthly job or incident-only recovery tool.
Inspect organization, repository and environment secrets, inherited process variables, named profiles, runner files and tool-specific configuration through approved access. Check the actual SDK or command-line credential precedence for the installed toolchain. The inventory is a set of references and ownership decisions, not a copied collection of secret values. Avoid downloading sensitive build artifacts into an uncontrolled folder to make the search easier.
Classify each consumer as required and ready to migrate, obsolete with owner approval, still dependent or unresolved. Retain the old source's relationship to each job until migration evidence replaces it. If one key is shared across unrelated accounts or applications, widen the change review to those owners before withdrawing it. A convenient repository cleanup must not silently disable another team's operational access.
3. Write the workload identity contract from observed claims
Owner: cloud identity administrator with workflow maintainer. Output: versioned identity contract. Record issuer, audience, subject shape and the intended execution context. Obtain controlled evidence from the job through an approved method without printing a reusable token. Keep enough non-secret identifiers to match the proposed conditions to the actual workload. Do not guess the subject from a repository URL and then broaden the pattern until authentication succeeds.
GitHub's AWS OIDC guidance describes audience and subject conditions, environment-specific subjects and immutable subject formats for applicable repositories. It also distinguishes requesting a token from resource-write permission. Use the format the repository actually emits. Our proposed acceptance record preserves that format and the configuration revision so future environment or repository changes can be evaluated against it.
Document workflow protections separately from identity strings. Which reviewed execution can reference the production environment? What prevents an unintended branch, altered workflow or untrusted input from running in the credentialed job? An environment name can appear in the accepted subject while its practical release controls are weak. Record those controls and their owners rather than assuming the cloud condition itself proves human approval.
4. Separate role trust from resource permissions
Owner: cloud identity administrator with security reviewer. Output: scoped role proposal. Specify who can assume the role and which operations the resulting session may perform. AWS's OIDC role guidance treats trust and permissions as separate policies and recommends restrictions on accepted GitHub subjects. Our review follows those separate layers, with explicit forbidden contexts and forbidden resource operations.
For the hypothetical release, name the artifact location, target service and intended environment. Avoid a role that can change every service simply because it is easier to migrate the current key's broad policy unchanged. Some operations may require additional resource or condition analysis. Retain those unresolved details instead of presenting a reduced policy document as proven least privilege before exercising the required workflow.
Choose an isolated role and harmless target for initial tests. Keep beta and production authority distinct even when the same action configures both. Pin reviewed executable dependencies according to the organization's supply-chain policy, and limit unrelated code in the credentialed job. A narrow subject cannot make arbitrary executable steps harmless once they possess the role's authority. These are application and execution concerns alongside the identity proposal.
5. Create the allow-and-deny test register before changing access
Owner: security reviewer. Output: test matrix and expected rejection layer. Define one accepted context and nearby forbidden contexts: an unintended branch, a different repository, a beta job requesting the production-scoped test role, an incorrect audience where testable, and an unintended resource action. Set the expected result and observation for each. Do not ask engineers to improvise negative tests after the successful path has already been accepted.
Keep token trust tests separate from permissions tests. A correct workload receiving a role but being denied a resource action tests the role's effective permissions. A rejected exchange tests identity acceptance only when the failure came from the intended condition. An unavailable endpoint, malformed request or missing provider setup is a test failure, not evidence that the trust boundary rejected an unauthorized identity correctly.
Specify harmless assertions that cannot damage shared resources. Verify role and target identity through an approved read or disposable fixture, then test the selected permitted behavior. Restrict negative tests to approved identities and targets. Do not manufacture a foreign organization's token or attempt access to an account outside the test authority. Unexercised cases remain unverified and need an alternate evidence method or explicit limitation.
6. Exercise replacement access with the old source unavailable
Owner: workflow maintainer with independent observer. Output: replacement execution record. Run one controlled consumer using the replacement identity. Remove the old source from that execution's approved configuration, checking inherited variables, profiles and runner state. Do not remove a shared secret globally while unrelated consumers remain unresolved. Record how the test excludes fallback and which identity the cloud operation actually used.
Capture run identifier, workflow revision, role, account, environment, target and observed service result with sensitive values removed. Authentication and deployment acceptance are different observations. A role can be correctly assumed while an upload fails, or a deploy step can report success against the wrong account. Have the service owner verify the actual fixture or artifact rather than accepting the command's exit code as the whole result.
Retain failed attempts and their reason. If a permissions gap is discovered, change the smallest reviewed boundary and repeat the affected tests. Do not reintroduce the old key automatically or grant broad administration to get a green run. Repeat the controlled replacement check for required scheduled and recovery consumers, or keep those consumers as an open migration condition instead of extrapolating from the primary release.
7. Prove denied contexts and recover from a failed exchange
Owner: security reviewer with recovery owner. Output: observed denial and recovery results. Execute the approved negative cases against the isolated role. Correlate each observed failure with the intended policy or workflow control. Retain provider error classification and restricted evidence references, without raw tokens. If a forbidden context receives access, halt progression and identify whether the defect is in subject matching, workflow eligibility or resource scope.
Rehearse an accepted job whose identity exchange fails for a controlled reason. Verify that the job stops before cloud mutation, reports a useful error and does not silently select the legacy credential. Restore the correct configuration through the approved recovery path and rerun the harmless check. An identity failure should not trigger uncontrolled retries, secret printing or a script that re-enables keys without an approval record.
Review what changed while the job was unavailable. A failed exchange does not mean no earlier deployment operation happened; an already running job may have acquired a session. Record run state and outstanding effects before retrying deployment. If a partial operation is uncertain, reconcile its service result first. The migration's recovery procedure should address operational state, not assume restarting authentication safely repeats every downstream write.
8. Cut over consumers under an owned observation plan
Owner: release approver with service owner. Output: cutover record. Approve the exact workflow and identity revisions, named consumers, observation period and recovery conditions. Stage consumers according to operational dependency rather than changing every job simultaneously. Keep the old path's temporary availability visible during the migration. Its presence is a remaining access risk, not a completed retirement state.
Observe required executions with the intended identity and accepted service outcome. Choose observation coverage from actual job schedules and criticality. A quiet day cannot demonstrate that a monthly integration works. Where a job cannot be safely exercised, document its alternative validation and remaining uncertainty. Do not close the organization-wide credential ticket when the approved scope covered only one repository's deployment jobs.
Hold progression if an unexpected credential source appears, a required job fails, or the operational result cannot be correlated with the run. A recovery decision must name whether it restores workflow code, permissions or another approved access path. Re-enabling the legacy key is a separate access decision with security review, not an invisible side effect of reverting a YAML file. Record any temporary exception and its withdrawal owner.
9. Disable the legacy key and read back its provider state
Owner: cloud identity administrator. Output: legacy-key state record. With required consumers and recovery accepted, withdraw the exact key identified in the change contract through the approved administrative process. AWS's access-key procedures distinguish deactivation, activation and deletion, and advise checking use before permanent deletion. Our proposed staged retirement keeps those actions separate so operational recovery and access closure remain inspectable.
Read back the provider-side state and record the identifier, time, operator and outcome. Removing a repository secret only removes that stored source. It does not establish the cloud key's state or eliminate copies elsewhere. If the provider readback is unavailable or ambiguous, keep the key withdrawal unverified. Do not treat a successful console click or a local configuration diff as the authoritative final state.
Where approved and safe, perform a harmless denial check for the withdrawn access path without exposing the secret. Use the organization's secure test mechanism rather than copying the key into a shell history. Continue observing required replacement consumers after withdrawal. If an unowned job fails, preserve that evidence and follow the recovery decision, including explicit approval for any access exception rather than quietly restoring the key.
10. Resolve existing sessions without assuming key withdrawal ended them
Owner: security reviewer with cloud identity administrator. Output: session-disposition decision. Identify session types the legacy access could have obtained, relevant roles, issue times and available activity evidence. Do not assume a disabled long-term key proves every previously issued session unusable. Establish applicable provider behavior and the closure method for the actual sessions. Record what remains unknown instead of making an instant-revocation promise.
AWS's role-session revocation guidance describes a deny policy for sessions issued before a cutoff, warns about impact to role users and documents exceptions. The procedure is not a universal per-key revocation button. Our recommendation is to review the role's other users and required operations before applying such a control, with an approved owner for any disruption.
Choose documented withdrawal, bounded expiry with adequate evidence, or another applicable approved control for the actual case. Distinguish future token exchanges from existing sessions and completed resource changes. Preserve outstanding operations and service state when selecting the cutoff. If evidence cannot identify affected sessions adequately, keep the exposure uncertainty open and obtain a security decision; do not conceal it behind the successful migration run.
11. Dispose of stored copies and retire obsolete recovery paths
Owner: workflow maintainer with security owner. Output: copy-disposition register. Revisit the locations inventoried before migration. Remove or restrict obsolete secret references according to their ownership and retention rules. Record runners, profiles, scoped secret stores and permitted artifact locations inspected. A broad claim that all copies are erased needs evidence beyond a repository search. State the coverage of the actual review and preserve unowned locations as open work.
Handle logs, backups and retained artifacts through the approved retention process. Do not delete shared archives or clear an entire runner home directory to eliminate one possible credential copy. Resolve exact targets and preserve records required for investigation. A retained copy whose provider-side access is closed may require restricted retention rather than immediate destruction; the security and data owners must decide the appropriate disposition.
Remove obsolete fallback behavior from workflow code and runbooks. Confirm that recovery instructions no longer ask an operator to reactivate the retired key routinely. Maintain the independently tested recovery path, with its owner and review trigger. If permanent key deletion is approved after the required evidence is accepted, perform that separate action and read back its outcome. This playbook itself does not authorize deleting a real credential.
12. Accept the evidence packet and define revalidation triggers
Owner: platform owner with security reviewer. Output: accepted migration record or owned exceptions. Reconcile every required consumer with its observed replacement result. Review the following acceptance criteria for denied contexts, resource scope, key state, session disposition and stored-copy coverage separately. Each incomplete item needs a responsible owner, next action and review trigger. Avoid an overall pass label that hides an untested disaster-recovery consumer or an unknown session lifetime.
- Required jobs ran with the intended role and accepted service result, without legacy fallback.
- Accepted claims and workflow protections match the reviewed configuration.
- Forbidden contexts and resource actions fail at their intended enforcement layer.
- Provider-side legacy-key withdrawal has an authoritative readback.
- Existing sessions have an approved and evidenced disposition, with limits stated.
- Stored copies and retained artifacts have scoped ownership and disposition records.
- Recovery access was exercised independently and cannot silently reactivate the legacy path.
- Unresolved consumers or evidence gaps remain visible with owners and follow-up.
Require revalidation when repository identity, subject configuration, environment controls, reusable workflows, runner images, credential resolution or role permissions change. Attach the evidence contract to the maintained configuration so a later edit can identify which acceptance result is no longer applicable. Use the supply-chain whitepaper for executable-artifact trust, and a cloud security review when the identity boundary or session closure remains uncertain. Reading and downloading require no email; contacting Ampity is optional.