Retiring Static Cloud Credentials Without Hiding the Risk
Design workload identity around narrow trust, controlled execution and resource permissions. Migrate with denial tests, refresh evidence and explicit retirement gates.
audience="Teams replacing stored cloud credentials in deployment and operational workloads." decision="Choose a managed runtime identity, external federation or a bounded exception, then prove which old access paths can be retired." position="Short-lived credentials are useful only within a controlled trust and permission chain. Accept the migration from observed allow/deny behavior, refresh and recovery tests, and evidence that fallback credentials are no longer usable." scope="A proposed engineering review, not a provider-specific configuration recipe or a description of Ampity's hosting. All workload records and test cases are illustrative." outputs={['A workload and credential register', 'A trust and resource-permission contract', 'An option comparison', 'An independent denial matrix', 'A refresh and outage rehearsal', 'An owned retirement decision']} />
Executive summary
Replacing a stored cloud key with temporary credentials removes one persistent secret from the ordinary execution path. It does not establish that the right workload obtained those credentials, that the executing code is trustworthy, or that the resulting permissions are appropriate. A narrowly scoped identity exchange can still hand an excessive role to compromised code. A correct role can still be bypassed by a forgotten key in the runner's environment.
This paper treats the change as an operational trust migration. Its position is to define the accepted workload, control who can change its execution, narrow the resources it can affect and prove retirement of the old path separately. The acceptance evidence includes rejected identities and resource operations, successful credential refresh, controlled dependency failure and a representative run after the old credential has been disabled. A green deployment is useful evidence, but it is not the entire acceptance decision.
The review applies to CI jobs, scheduled workers and service processes that access cloud APIs. It compares a platform-managed runtime identity, external workload federation and a time-bounded credential exception. These alternatives do not share identical claim formats, refresh behavior or revocation semantics. Choose from the actual deployment constraints rather than forcing every workload into the same OIDC tutorial.
The proposed diagrams describe generic enforcement boundaries, not a deployed cloud topology. The examples use fabricated workloads and inert test operations. Official provider documentation supports specific authentication, policy and lifecycle distinctions; the migration gates and evidence method are engineering recommendations. This is not a security certification, a guarantee against credential theft or permission to change an organization's production identity configuration without its normal approval.
Define what is changing and what is not
A static credential is a reusable secret or key material distributed to a workload outside a short-lived runtime issuance flow. Workload identity here means using an asserted or platform-attached identity to obtain access appropriate to that workload. The target state still uses credentials or tokens at runtime. Calling it passwordless must not imply that nothing sensitive exists, that credentials cannot be stolen or that every access path has disappeared.
Define the unit of migration. A repository is often too broad: build, test, preview deployment, production deployment and maintenance jobs can require different resources and approvals. Likewise, a service account name is not a complete description of the processes that run under it. Separate workloads when their accepted identities, privileges or operators differ. Record the release, runtime and maintenance paths rather than assuming the busiest job represents the rest.
State the unchanged business requirement. A deployment still needs to publish an identified artifact to the intended target. A reconciliation worker still needs access to the permitted records. Identity migration must not silently broaden the data set, move approval to a different team or replace a transactional operation with an uncontrolled script. Use the existing service acceptance criteria alongside the identity-specific criteria.
Identify deliberate exclusions. Third-party systems that only accept stored API keys may remain outside this migration. Human emergency administration can need a different access model. A legacy appliance might not support token exchange. Record each exception, its owner and its review trigger. Excluding a workload can be reasonable; presenting an exception as migrated because its secret moved to another vault is misleading.
Inventory execution paths before choosing a provider mechanism
Start from the places credentials are consumed, not just the secret store. Inspect runner configuration, environment variables, shared credential files, container images, deployment templates, helper scripts and scheduled processes. Include credential injection outside the visible workflow, such as a runner service or organization-level integration. Avoid printing secret values during this work. The useful inventory contains references, ownership and permitted use, not another copy of credentials.
For each workload, record the source of identity, executable configuration owner, target account or project, required operations and normal invocation contexts. Include manual reruns, branch builds, environment deployments and disaster-recovery procedures when they are supported. A monthly maintenance job may use the old credential even after daily deployments have switched successfully. A quiet observation interval is not proof that the credential has no remaining consumer.
Collect the actual credential source selected by the application's SDK or tool in a controlled test. AWS's standardized credential-provider documentation explains that provider chains and supported mechanisms vary across SDKs and tools. Do not assume that adding federation configuration makes an existing environment credential irrelevant. Verify the specific client, version and process environment that will run the workload.
Keep the inventory connected to evidence. A workload marked migrated should identify the reviewed configuration and test run that established its selected identity. Unknown consumers remain unknown until investigated or explicitly covered by a retirement test. Secret-store deletion alone does not establish that copies in cached files, old runner disks or retained artifacts are unusable. Retirement needs a provider-side disposition as well as local cleanup.
Separate issuance, trust and resource authorization
The identity issuer asserts properties about a workload. The cloud's federation or runtime mechanism decides whether to accept that identity. Resource authorization determines which operations the resulting principal may perform. The workload's executable control determines what the trusted process actually does. These boundaries interact, but no single policy file settles all four questions.
AWS's OIDC role guide separates the role's trust configuration from its permissions policies and recommends restricting GitHub subjects and protecting deployment environments. A successful role assumption proves that the exchange accepted the presented identity under the observed configuration. It does not establish that every subsequent service action is permitted or that the accepted workload should have been trusted.
AWS's web-identity exchange reference documents temporary credentials and the relationship between role and session policies. Session policies restrict rather than expand the role's identity-policy permissions. Other applicable policy layers and resource conditions still matter. This paper does not reduce effective authorization to one policy intersection or suggest that an issued session bypasses resource-level enforcement.
For the review contract, write separate intended outcomes for identity exchange and resource access. An unintended branch should not obtain the deployment role. An accepted deployment session should not read an unrelated private data set. A failure at the first boundary does not test the second. Preserve the reason for denial so a broken network connection or expired test token is not mistakenly counted as policy enforcement.
Map the trust chain the workload will actually use
The first visual asks which independent boundaries stand between executable workload code and a resource operation. It intentionally uses generic nodes because the same question can apply to a managed runtime or external federation. The details of issuer, audience, subject, token exchange and target principal belong in the accompanying register for the selected mechanism.
The main path shows issuance, trust admission and an authorized resource request. The separate code-control boundary matters because a valid identity can be used by changed executable steps. The rejection paths mean no permission is established at that boundary; they do not imply that a denied request reverses earlier effects. This is a proposed review model, not evidence of any deployment's implemented controls.
For each connection, name the authority and the observable record. The runner can identify its execution context; the issuer can attest supported claims; the cloud can record an exchange; the resource can record an allowed or denied operation. Avoid substituting a human-friendly job label for a claim the provider evaluates. Keep sensitive tokens out of the register and retain only the sanitized fields necessary to reproduce the review.
Also identify links that depend on implicit behavior. A credential helper might refresh automatically, an environment rule might select an approval requirement and a reusable workflow might hide a role exchange. Resolve those dependencies before declaring the chain controlled. A diagram is a useful conversation aid only when its arrows correspond to actual mechanisms and its unknowns remain visible.
Inspect identity claims rather than copying example strings
GitHub's AWS OIDC guidance distinguishes audience and subject conditions, environment-based identity and token-request permission. It notes that subject formats can differ for repositories using immutable identity claims. Verify the emitted format for the selected job rather than copying an older subject pattern. The paper deliberately avoids providing a ready-to-apply trust policy with fabricated organization identifiers.
Define which claims the cloud provider can enforce in the actual integration. A claim visible in a token is not necessarily a supported policy condition at the destination. Record the intended issuer, audience and workload restriction with the provider's enforcement mechanism. If a desired restriction cannot be expressed there, place an evidenced control elsewhere or narrow the supported execution path. Do not describe the unexpressed restriction as enforced.
Review identity lifecycle changes. Renaming an environment, changing reusable workflow configuration or transferring repository ownership may alter the relationship between a friendly name and the accepted principal. Decide which changes invalidate the migration evidence and require review. Prefer stable identifiers where the integration supports them, but verify their behavior rather than assuming all provider subject formats are immutable.
Inspect both legitimate and unintended contexts. A token can be correctly signed by the expected issuer yet belong to the wrong workload. A valid audience does not identify the right repository. An exact repository match does not necessarily distinguish production from preview execution. Construct test cases from the business restriction and the actual emitted identity, with sensitive assertions retained only in the approved test environment.
Choose between runtime identity, federation and an exception
A managed runtime identity can be a good fit when the workload already runs on a platform that attaches identity and supplies supported credentials. Its advantage is less external identity-exchange configuration in the workload. Its trade-off is that runtime placement, attachment permissions and process isolation become important control surfaces. Any code able to use the attached identity must fall within the accepted execution boundary.
External federation is often appropriate when CI or another external runtime needs cloud access. It can replace a distributed cloud key with a trust relationship and temporary access. Its trade-off is dependence on issuer availability, claim semantics, token exchange and compatible client behavior. It still needs a narrow principal and a controlled workflow. Centralizing trust can simplify management while increasing the consequence of a misconfigured shared admission rule.
A bounded static-credential exception may be necessary where neither mechanism is supported. It should have a narrower purpose, owned rotation and withdrawal procedures, and an expiry or review trigger. Do not leave a general production key available to all jobs merely as insurance against federation errors. An exception is a consciously accepted residual risk, not an automatic recovery mechanism for the entire migration.
Compare these choices against the same workload requirements. Include resource scope, execution isolation, refresh needs, outage behavior, auditability, support burden and retirement evidence. A mechanism with the fewest workflow lines can still be harder to operate if it hides credential selection. A more complex exchange can be appropriate when its identity and recovery boundaries are explicit. Choose the smallest access path that meets the requirement, not the shortest configuration snippet.
Preserve provider differences in a multi-cloud decision
Google Cloud's federation overview distinguishes attribute mapping, admission conditions and access through direct principals or service-account impersonation. It also describes limitations affecting some APIs. Select the access path for the actual target API and record both the identity admission rule and the resource grant. A pool accepting an identity does not independently establish permission to every resource in a project.
Microsoft's workload identity federation overview describes external identity trust for supported workloads and token exchange without maintaining the corresponding secret in the workload. The applicable identity type, federated credential configuration and downstream resource grants need their own review. Do not translate an AWS role condition into an Entra configuration by renaming the fields and assume equivalent semantics.
Use a provider-specific appendix or register when a workload crosses several clouds. Keep the original principal, mapped identity, target role or account, selected resource policy and supported withdrawal procedure visible. Additional impersonation or role-chaining steps can introduce another trust decision and another audit join. A multi-cloud diagram that shows one generic Login arrow conceals precisely the distinctions the migration needs to test.
Provider abstraction has limits. It can standardize the application's request for credentials and the shape of an evidence record, but it cannot make all lifetimes, claim conditions, propagation delays and denial reasons identical. Document where behavior differs and provide a test for the chosen mechanism. Operational consistency comes from consistent review questions and ownership, not from pretending different providers have interchangeable enforcement.
Control the code behind the accepted identity
An identity rule identifies a workload context, not the safety of every command running within it. Review who can edit the workflow, modify dependencies, publish a referenced action or change runner configuration. Connect production permission to the artifact and executable steps the owner intended. A legitimate release identity running newly substituted code can satisfy the trust rule while violating the deployment's purpose.
Separate jobs that need production access from those that process untrusted contributions. Identify when the execution reads repository-controlled scripts, downloaded build tools or generated configuration. Use the platform's supported protections, review and pinning procedures appropriate to those dependencies. This paper does not claim that a protected environment alone makes arbitrary code safe or that one pinned dependency establishes the integrity of the entire build.
Runner reuse deserves its own review. A previous job might leave credentials, mutable files or a credential-helper configuration that affects a later job. Determine the isolation contract and the actual cleanup evidence. A self-hosted runner with production reachability needs controls for both identity issuance and the operating environment. Temporary credentials can still be copied during their validity if the process boundary is compromised.
For AI-assisted deployment work, keep model-generated scripts and recommendations outside the permission decision. AI can help prepare a diff or investigate a failure, but a suggestion to broaden a trust wildcard is not an approved security change. Require the same artifact review and independently defined denial tests used for human changes. Automation speed does not replace evidence that the accepted job still performs the intended work.
Establish a scoped migration pilot
Choose one representative workload with an understood owner and a bounded target. Prefer a harmless test resource for initial exchanges and denial cases. Define its expected identity and required resource operations before enabling the new path. The pilot should exercise more than the easiest success case, but it should not intentionally attempt destructive access to production data to prove a restriction.
Keep the old and new paths distinguishable during the pilot. If they coexist temporarily, record which path each run selected and why coexistence is allowed. Do not combine their credentials in one opaque provider chain and infer migration success from a completed job. A pilot that silently falls back to the old key has not established the viability of the new path.
Set progression and pause criteria owned by both operations and security. Examples include correct principal selection, required harmless action success, rejected unintended identities, rejected out-of-scope actions and refresh without a stored key. These are proposed criteria, not universal thresholds. The service owner should also confirm the deployment's ordinary outcome, latency and recovery requirements under representative conditions.
Record configuration revisions and test identity separately from sensitive tokens. Retain enough evidence to explain why the pilot passed after the workflow changes later. If the issuer or SDK behavior differs between the pilot and target production environment, treat that as an untested boundary. Approval to run the pilot is not approval to copy its broad test permissions into production.
Build a denial matrix with independent expected outcomes
The fabricated workload Release-C publishes only a selected artifact to target T1. Preview-C must not obtain that production-scoped identity. Release-C must not read an unrelated data collection T2. The test matrix separates valid token admission from allowed resource work and from runtime correctness. An independently specified expected outcome prevents the implementation from defining success merely by returning its own completed flag.
| Controlled case | Expected boundary | Required evidence | | --- | --- | --- | | Release-C with the declared execution identity | Admit the selected principal | Sanitized identity fields and actual resulting principal | | Preview-C presenting a valid issuer token | Reject production admission | Relevant condition failure, not an unrelated transport error | | Release-C requesting an out-of-scope resource | Reject the resource operation | Selected principal and harmless denied request | | Release-C after the static key is disabled | Complete via the new path only | Credential-source record and representative service outcome | | Release-C interrupted during a write | Preserve uncertain effect state | Original operation reference and destination readback |
Run denial cases with real enforcement in an approved test scope, not only a local policy mock. Static analysis can identify an obvious wildcard, but it cannot prove the exact issuer and client interaction works. Conversely, a failed exchange caused by an invalid setup does not prove the intended subject restriction. Record why the failure occurred and repeat a controlled valid token for the intended rejected context.
Include a negative control in an inert harness to show the observer notices overly broad admission or hidden fallback. Do not weaken production policy for this control. The observer should distinguish admission, selected principal, resource decision and resulting effect. Missing telemetry is an evidence gap, not a pass. Preserve rejected cases as well as the accepted case in the pilot's decision record.
Prove refresh across the workload's actual lifecycle
A short test can finish before the first credential expiry and conceal a refresh failure. Exercise a representative longer-running operation, scheduled continuation or delayed retry where the client must obtain fresh credentials. Confirm the selected credential source before and after refresh. The useful result is not that a token has an expiry field, but that the workload behaves correctly when its credentials need renewal.
Separate issuer assertions, exchanged cloud credentials and application operation lifetimes. Their validity windows can differ. A process can retain issued credentials after the issuer token used for exchange expires. A retry may occur after the job context that originally obtained them has ended. Define which refresh is supported and which continuation should instead stop or request a newly authorized execution.
Do not solve refresh errors by embedding a permanent key in a helper. That creates a second path with a different trust model. Investigate the actual SDK behavior, token source, permissions and runtime availability. A retry budget should include issuer and exchange calls so several layers do not amplify an outage. Limit retry work by the operational deadline and the consequence of repeating the resource action.
Distinguish authentication retries from business-operation retries. Refreshing credentials can restore the ability to query a destination, but it does not prove a previous write failed. Preserve an operation identifier and reconcile destination state before repeating an uncertain change. A migration that improves identity but introduces duplicate deployments or repeated external effects has not met the workload's original acceptance criteria.
Rehearse issuer and exchange failure without broadening trust
The new dependency chain may fail at issuance, exchange or resource authorization. Inject controlled failures and identify the resulting state. The workload should expose whether no credential was obtained, a principal was obtained but access was denied, or a resource operation was already issued with no receipt. These states require different recovery decisions and should not be compressed into one login error.
Use bounded retries for transient dependency errors where permitted. Check the job's deadline and downstream capacity before resuming backlog. A security denial caused by changed claims should not be retried indefinitely as if it were an endpoint outage. A successful probe does not establish that all queued jobs still have current permission or that all partially completed writes can be restarted safely.
Define an owned degraded mode. It may hold deployments, permit a narrower diagnostic action or invoke an explicitly approved emergency process. It should not widen the accepted workload set or automatically re-enable a general static key. Communicate what is unavailable and what remains safe. The responsible operator needs a reason and a bounded next action, not a suggestion to make the trust policy more permissive until the pipeline turns green.
Preserve evidence through recovery. Record which configuration was active when the failed attempt ran, whether a session was issued and whether an external operation may exist. Resume with the original operation identity where appropriate. If readback is incomplete, retain uncertainty and stop before another consequential write. A restored identity service solves access availability, not the accounting of effects already attempted.
Distinguish withdrawal from the end of an issued session
Removing permission to obtain a new session and constraining an existing session are different operations. AWS's role-session withdrawal documentation describes a specific policy-based mechanism for sessions issued before a selected time, with important applicability conditions. Review the mechanism for the actual credential type and permission scope. Do not promise instant universal revocation merely because a trust relationship was changed.
Record the incident questions separately: can new credentials still be issued, can existing credentials still call the target, and what actions already occurred? The last question requires destination or activity evidence. Denying a future request does not undo an earlier deployment or disclosure. An emergency procedure should explicitly include investigation and corrective actions owned by the affected service, not only identity administration.
A withdrawal rehearsal should use approved test sessions and harmless operations. Compare an already issued session with a fresh exchange after the change. Capture the reason and timing of each denial, and note policy propagation or observation limits. If one mechanism does not cover another session type, identify the gap and its separate control. An attractive single status labelled Revoked should not conceal those distinctions.
Keep rollback and incident response separate. Rollback addresses an incompatible migration configuration. Incident response addresses suspected compromise and may make the former credential path unacceptable. Re-enabling that path during an incident can restore the attack opportunity. The authorized owner needs to choose a bounded recovery route from the current risk, not mechanically reverse the last configuration change.
Retire the static path as a separate acceptance decision
Retirement requires more than observing the new path succeed once. Locate remaining consumers, disable the credential at its authority, remove approved local references and exercise representative workloads without it. Keep the credential identifier, owner and resulting disposition in the evidence record without storing the secret. A deleted pipeline secret can leave a copied key valid elsewhere; provider-side withdrawal matters.
Choose the observation window from actual workload schedules and recovery procedures. Include infrequent maintenance or seasonal jobs through controlled rehearsal when waiting is impractical. Zero recent activity does not prove no dependency. If coverage is incomplete, retain the exception and its risk explicitly instead of calling the migration complete. The owner may accept a staged retirement, but the acceptance record must state which consumers remain.
The second visual asks which evidence permits progression from pilot to retirement and where a failure returns the work for investigation. This is a decision lifecycle, not another runtime architecture. A failed gate does not authorize recreating the old key automatically. The controlled exception, if needed, has a separate owner and conditions outside the ordinary path.
After retirement, remove the fallback selection logic and verify that a broken exchange fails visibly rather than finding an old credential source. Preserve deployment and recovery documentation for the new path. A migration is not complete when operators must remember to choose the safe setting by hand every time. The default behavior should match the approved contract, with any exceptional access deliberately separated.
Use a review checklist with explicit ownership
The workload owner should confirm the required service outcome, supported execution contexts and recovery deadline. The security owner should confirm accepted identities, rejected identities, resource scope and session-withdrawal limits. The platform maintainer should confirm credential selection, runtime isolation, refresh and failure behavior. These responsibilities can belong to one person in a small team, but their evidence questions remain different.
The decision record should contain the workload register, chosen mechanism and alternatives rejected; sanitized claim and principal evidence; reviewed policy and workflow revisions; allowed and denied tests; refresh and dependency-failure observations; uncertain-effect handling; credential disposition; consumer coverage; and open exceptions. Require an owner for each unresolved item. A checklist tick without the referenced observation should not be treated as acceptance evidence.
Revisit the decision when the issuer, repository identity, environment protection, runner isolation, SDK, reusable workflow, target permission or resource scope changes. These triggers are more useful than a promise to review annually while operational changes occur every week. Store the test cases alongside the controlled implementation and retain a meaningful way to reproduce them without access to old secret material.
Define the final acceptance narrowly and honestly. For example, the reviewed workload can deploy the selected artifact to T1 through the new principal, selected unintended contexts and actions are denied, refresh and bounded failure handling have been exercised, and the recorded static credential is disabled with representative consumer coverage. This statement does not certify every cloud workload, rule out all compromise or claim that all exceptional access has been removed.
Limitations and a practical next step
This review does not replace threat modelling, provider-specific policy analysis or runtime assessment. Unsupported exchanges, claim restrictions or resource grants remain gaps. A managed identity cannot isolate credentials from unrelated code sharing its execution context. Narrow the execution or record an explicit exception.
Test provider availability, propagation timing and deployment effects in the approved environment. Local fixtures do not prove rejection of a real token or effective session withdrawal. Distinguish configuration review, controlled enforcement evidence and operational coverage.
Begin with one workload and its credential register. Define an unintended identity and resource operation that must be denied, then design pilot, refresh and retirement tests around them. Resolve unknown credential selection before counting a successful deployment as migration proof.
Continue with the static CI credential replacement playbook and credential retirement evidence article. Ampity's cloud security work and DevOps and SRE services can help scope the migration. Reading and downloading the PDF require no contact details; enquiries remain optional.