Decide What May Run Again in a Step Functions Redrive

Bind one Standard execution to a state-visit manifest, original definition and input, repeatable effects and current authority before a bounded redrive rehearsal.

Scope: one execution, not permission to retry every effect

Prepare the state-visit manifest before using execution redrive. A failed workflow can contain an already accepted external action. A previously successful branch can also become repeat work under a different enclosing-state error. The recovery question is not simply whether the console offers a button. It is whether this particular repeat path remains both explainable and permitted.

This playbook supplies an unexecuted educational rehearsal for one Standard execution in one reviewed sandbox account and Region. Task fixtures use existing inert request-response integrations with an inspectable operation ledger. Parallel and Inline Map visits are in scope. Top-level Express, Distributed Map, nested workflows, callback tokens, .sync jobs, production recovery and definition/input migration are not runnable scope. Discovering one of those constructs means HOLD and a separate procedure, not skipping its rows.

AWS documentation checks were refreshed October 8, 2026. This is an educational rehearsal plan, not a record of an executed AWS or customer recovery. The event-driven systems playbook owns outbox, duplicate guards and general business replay. This procedure owns the narrower relationship between an execution's state history and its possible redrive effects.

1. Bind the execution and separate the actors

Owner: evidence reader and security reviewer. Output: scoped identity packet. Record account, Region, execution ARN, state-machine identity and workflow type from authorized evidence. A friendly execution name is not sufficient across accounts or Regions. Record the application operation identity separately. The execution ARN is not the business deduplication key by default.

Use the Step Functions authorization reference to resolve exact API callers and resource scope. Separate the reader's history/definition access, the operator's states:RedriveExecution, the independent stop operator's states:StopExecution and the state-machine execution role's downstream permissions. Record applicable organizational and resource-policy restrictions. The ability to operate a workflow does not confer authority to submit another supplier change or notification.

Agree permitted evidence handling before exporting history. Inputs, causes, output and resource identifiers may reveal sensitive information. Store exact data in the authorized repository; share references, hashes and synthetic aliases in broader reviews. KMS or payload restrictions are not reasons to substitute an administrator credential or silently widen a policy.

Gate: wrong identity, unsupported workflow type, unresolved evidence permissions or an unverified stop owner means HOLD. Do not create a new execution to evade the gate.

2. Check current eligibility without treating it as authority

Owner: workflow maintainer. Output: dated eligibility row. The RedriveExecution API applies to unsuccessful Standard executions: failed, aborted or timed out. Its documented boundaries include the 14-day redrivable period from completion, maximum open time of one year and history count below 24,999. The redrive guide additionally requires an original start on or after November 15, 2023.

Capture status, dates, redriveStatus, redriveStatusReason, count and last redrive date through DescribeExecution. RUNNING and SUCCEEDED are not candidates for this action. Record check time and source. This read is eventually consistent, so a stale REDRIVABLE record is not an immutable grant. Recheck at the approved execution boundary and treat contradictory or unavailable evidence as unresolved.

Keep a future rehearsal's end time separate from the provider eligibility window. Do not plan another attempt at the last instant of a nominal date range. A workflow record cannot renew an expired business instruction. Likewise, a fresh business approval cannot remove an AWS eligibility restriction.

Gate: a provider-ineligible execution is rejected for redrive; incomplete eligibility evidence is held. A separately designed new execution may be a later decision, but this playbook does not authorize it.

3. Recover the original basis and inventory mutable dependencies

Owner: workflow maintainer and task owner. Output: original-basis manifest. Preserve the execution-associated definition, input reference, revision and originally associated version/alias. Use DescribeStateMachineForExecution, not only today's state-machine definition. Metadata-only access to an encrypted definition does not supply the missing ASL. Record unavailable fields explicitly.

AWS redrive retains the original execution ARN, definition and input. If the original attempt used a version or alias, moving that alias does not make redrive use its newer workflow version. A changed definition requires a separate new-execution decision, not an assertion that redrive picked up the patch. The API request accepts an execution ARN and optional client token, not a replacement business payload.

Now inspect what that ASL references. A Task may reference a mutable function alias, an unqualified resource, current service configuration or a destination whose authorization has changed. Original ASL identity is not proof of unchanged code or external state. Record the concrete task implementation and configuration basis available to the proposed attempt. If an original dependency cannot be reconstructed, state what is unknown instead of inventing a historical version.

Gate: if recovery needs a changed payload or workflow definition, hold this procedure and hand off the changed instruction. If task basis changed, the application owner must assess repeatability under that actual basis before proceeding.

4. Build a visit-level history, not a list of state names

Owner: evidence reader and workflow maintainer. Output: complete visit inventory. Obtain every page through GetExecutionHistory, keeping request arguments consistent while following tokens. Tokens expire after 24 hours. A single latest-events page cannot prove the complete execution or history-event count. Record page completeness and capture times.

Identify each relevant visit by nested state path, branch or iteration identity, entry/exit events and original disposition. The same state name can appear in multiple branches or repeated visits. Link attempts to their particular visit rather than merging them into one successful badge. If execution data was omitted or truncated and a required input is unavailable, preserve UNKNOWN and request the permitted evidence, not a guessed payload.

Inventory three categories: successful work potentially preserved, unsuccessful work potentially reentered, and downstream work that may execute for the first time after recovery. The third category matters: a redrive can reach a notification or mutation that the failed attempt never reached. Absence from prior history is not an exemption from current authority and effect review.

Gate: incomplete state paths, missing history pages or ambiguous effect ownership hold the manifest. History says what orchestration recorded, not necessarily what a destination applied before a lost response.

Filled fictional pair: identical Task names, different visits

This pair is an unexecuted educational packet, not an AWS response or destination receipt. The original ASL hypothetically names a Task Apply in each of two branches of ReviewParallel. Paths are analyst-derived from that ASL and linked history, not a provider-supplied fullPath field. Event references below are stipulated numbers inside this fictional history only. They must never be copied into an actual execution record.

Visit A: path ReviewParallel / branch[0] / Apply / visit[0]; fictional entry event 8, successful Task result event 12 and state-exit event 13. Stable inert operation identity fixture-A-apply, request fingerprint fixture-A-v1; prior destination disposition CONFIRMED EXISTING under fictional receipt ledger-A-1. Current authority EXPIRED; execution-time guard must refuse a new effect. Enclosing ReviewParallel failure event 29 is stipulated States.DataLimitExceeded. Proposed classification: potentially reentered despite prior success, but effect admission HOLD. A prior successful authorization or receipt does not renew current permission.

Visit B: path ReviewParallel / branch[1] / Apply / visit[0]; fictional entry event 18 and Task-failure event 22, with no established successful exit event. Stable inert operation identity fixture-B-apply, request fingerprint fixture-B-v1; prior destination disposition UNKNOWN because the fictional destination receipt is missing. Current authority GRANTED within the stipulated window, with an existing inspected guard, but this cannot resolve whether the prior action occurred. Same enclosing failure event 29. Proposed classification: potentially reentered failed visit; effect admission HOLD until destination readback resolves the original operation identity, not a newly minted identity.

For the same child outcomes under an ordinary enclosing error instead, Visit A would be preserved and Visit B potentially reentered. That alternative does not change B's unknown-effect hold. Under the stipulated data-limit condition both visits belong in the repeatability review, and this whole packet remains held. The pair demonstrates why grouping by the name Apply would conceal different history, operation identities, authority and dispositions. No redrive, destination query or guard test was executed for this example.

5. Expand the reentry set from the exact enclosing error

Owner: workflow maintainer. Output: justified reentry classification per visit. Apply the individual-state redrive behavior to the observed enclosing-state outcome. Ordinary unsuccessful Task work is rescheduled. Choice rules may be reevaluated, a past timestamp Wait may advance and a Fail state may fail again. Do not promise progress merely because another attempt is possible.

For a Parallel or Inline Map state, ordinary redrive selects failed or aborted branches/iterations. If that state failed with States.DataLimitExceeded, the whole state reruns, including work that succeeded before. This is why a successful child row cannot be excluded based only on its local badge. Identify the enclosing failure and its effect on descendants.

In a stipulated Parallel state, an ordinary failure preserves the completed A branch and reenters failed B. A DataLimitExceeded enclosing failure reenters both, including previously completed A. These are potential state visits, not observed effects.

*Illustrative comparison of one Parallel state. Gray records preserved work; blue records reentry candidates. It does not imply authorization, observed account behavior or that an external action is safe to repeat. Inline Map needs the corresponding iteration-level inventory, not a branch label copied into its record.*

The current error-handling reference allows an explicit Catch or Retry match for States.DataLimitExceeded; States.ALL does not catch it. Do not repeat the obsolete claim that this error is universally uncatchable. Record what the original ASL actually does. A handler or data-shaping change written after failure is not inserted into the preserved definition by redrive.

Gate: unexplained error, exception scope or descendant path means HOLD. Redrive does not fix an unchanged oversized input/output condition by itself.

6. Reject a different Map or integration procedure explicitly

Owner: workflow maintainer. Output: construct applicability decision. Distributed Map redrive is governed through the parent workflow and Map Run, not the Inline Map rule above. Standard children resume unsuccessful state paths where individually eligible. Express children restart at their first ItemProcessor state using StartExecution. Data-limit/runtime failures can start a new Map Run, and ItemReader input behavior depends on how far the original run progressed.

Therefore a Distributed Map is a named handoff, even if its parent is Standard. Do not infer that every child has the parent's eligibility window or that every S3 source reread preserves original input. Existing Map Run operations may still be finishing after a parent stops. This procedure does not supply its IAM, data-source or child-attempt controls.

Likewise, integration patterns distinguish request-response, .sync and callback work. A task ARN suffix changes the operational contract. Exclude nested orchestration, .sync and callback fixtures from this rehearsal; name an owner for any such path found in the actual inventory. Broadening scope requires a new review, not deletion of awkward rows.

Gate: every potentially reached state must be within the admitted fixture contract. Unsupported or unknown integration behavior stops admission for the whole selected execution, not only that row.

7. Reconcile effects and recheck current business permission

Owner: application and business-action owners. Output: effect disposition for every possible repeat/new action. Join each Task visit to its stable operation identity, request fingerprint, destination, actual prior receipt and current authority. Distinguish no effect, confirmed existing effect, authoritative refusal and unresolved outcome. A task timeout does not prove no effect occurred.

Use the business reconciliation article for that domain review. Do not copy SQS transport-identifier behavior into Step Functions execution identity. The principle carried forward is narrower: resolve effect evidence and permission at the owning boundary before repeating a consequential action.

A successful authorization state earlier in the workflow can remain preserved. Redrive is therefore not a promise that the original approval will be checked again. Verify that an already implemented guard at each relevant execution boundary consults the required current authority and preserves original effect identity. If permission can expire after a preflight review and no execution-time guard exists, HOLD. A signed spreadsheet is not a substitute for the missing guard.

Check guard retention against the full proposed recovery period. A destination's finite deduplication window or deleted operation record may make an old action unsafe to repeat. For already applied work, the expected result might be a confirmed no-op, not another write. For unknown work, retain reconciliation HOLD. Do not invent a new operation ID merely to force acceptance.

Gate: any repeat/new effect lacking supported current authority or prior-outcome evidence holds the selected execution. Workflow eligibility, retry budget and IAM permission cannot override this gate.

8. Bound refreshed retries as calls, effects and time

Owner: workflow maintainer, task owner and cost approver. Output: attempt/exposure worksheet. Retry behavior resets defined retry counts to zero for reentered Task, Parallel and Inline Map states. MaxAttempts is the retry count, excluding the initial call. Inspect matching error names, retrier ordering, loops, parent retries and nested state repetition before producing a maximum.

For one hypothetical leaf Task, assume exactly one matching retrier with MaxAttempts: 2, no enclosing retries or loops, and one permitted redrive. The original visit can issue one initial call plus two retries, or three calls. Redrive can admit another three, so the combined ceiling is six calls under those assumptions, not two or four. Six calls also does not mean six business effects: a supported identity guard can return the already existing operation. This is arithmetic, not an executed AWS test or a recommended retry setting.

Do not multiply every state by three. A retry on an enclosing state can repeat several descendants; a loop can revisit a leaf. If the maintainer cannot bound the actual path, record exposure UNKNOWN and hold the exercise. Independently account for state transitions, task invocations, logging/storage and downstream charges. The cost owner supplies current Region/service prices and a ceiling before any authorized run. No cost was incurred or estimated for this draft.

AWS resets defined state-machine and Task timeout clocks on redrive as documented. Our required control is a separate wall-clock rehearsal deadline and current business expiry. Do not convert a refreshed technical timeout into permission to continue after the action grant or spending envelope expires.

Gate: unknown call exposure, absent approved cost ceiling or an expiry that the actual guard cannot enforce means HOLD.

9. Specify independent fixtures before changing anything

Owner: independent observer with application owner. Output: frozen expectations. Use existing inert fixtures and inspectable operation records. The table supplies expected dispositions to challenge a candidate; it is not an account result or a universal state simulator. Prepare distinct approved executions for changed enclosing-error cases rather than pretending to alter an original execution's error by editing its history.

FixtureExpected admission or observed assertionEvidence to collect
Completed predecessor; failed TaskPrior completed work preserved; unsuccessful Task reenteredOriginal/added event paths and effect identity
Parallel A completed, B failed; ordinary errorA preserved; B reenteredBranch-specific visits and destination receipts
Same branch statuses; enclosing DataLimitExceededA and B potentially reenteredExact enclosing error and repeated branch visits
Inline Map iteration 0 completed, 1 failedOrdinary error preserves 0; data-limit failure can repeat bothIteration identifiers, not aggregate Map badge
Exhausted simple two-retry leafNew budget permits up to three additional calls under section 8 assumptionsAttempt history versus committed-effect count
Effect applied but receipt lostHOLD until prior effect reconciled; no fresh operation identityOriginal-operation destination lookup
Authority expired, AWS eligibility positiveHOLD, despite technical eligibilityCurrent guard decision and no unauthorized effect
Workflow alias movedOriginal workflow version still governs redriveAssociated version plus actual mutable task basis
History or encrypted definition unavailableUNKNOWN and HOLDMissing page/field and named evidence owner
Distributed Map, Express or unsupported integration discoveredSeparate handoff, no fallback mutationType/resource evidence and receiving owner

A local validator can check these stipulated classifications and arithmetic. It cannot prove AWS transition semantics, code behavior, guard enforcement or effect counts. Actual rehearsal acceptance requires those observed records, including inconclusive fixtures and failures in the denominator.

Gate: do not repair the candidate after seeing a failing fixture and relabel the old run passed. Record the revised candidate and a separately authorized rerun.

10. Authorize one mutation and preserve its request identity

Owner: business-action owner and authorized operator. Output: one bounded execution decision. Review the complete manifest, eligible state, actual dependency basis, repeatable effects, fixture expectations, deadline, spending ceiling and stop path. Authorization is for the identified manifest revision and effect envelope, not all executions matching a state-machine prefix.

Execution evidence, each possible effect's current authority and the funded stop envelope must all support a specific redrive manifest. A separately approved mutation then requires history and destination reconciliation. Missing evidence holds admission, not an automatic retry.

*Proposed evidence and authority boundary, not a runtime service topology. Blue routes carry reviewed evidence. Brown means missing evidence holds the decision. A redrive response is not the destination's business settlement, and no mutation was performed for this reference.*

The redrive API token contract provides bounded request idempotency. Preserve the submitted token, operator, request time and response/readback reference. If the response is lost, inspect permitted execution evidence before repeating the request. A fresh token can represent another mutation, and token protection never substitutes for downstream business deduplication.

This article intentionally gives no executable redrive or StartExecution command. A reviewer must bind the actual tool/API request to the approved execution and retain a separately verified stop route before acting. No local script in this package contacts AWS.

Gate: changed manifest, expired authority, inconclusive request result or mismatched ARN stops further admission. An operator cannot renew their own grant merely because the first attempt failed.

11. Observe the attempt and reconcile the destination

Owner: independent observer and application owner. Output: joined result evidence. Capture the appended ExecutionRedriven event, redrive count/date and state-visit history, then compare each expected visit with its observed disposition. Keep observation times and complete pages. Do not infer a missing new event is absent forever from one eventually consistent read.

Join orchestration results to the destination's operation records. A request-response task can progress after a service response without proving every downstream business consequence. Define the actual receipt contract of the inert fixture before evaluating it. A successful execution is not sufficient when the intended disposition was suppression of a duplicate or refusal of expired work.

Record three independent outcomes: request accepted/not established, orchestration result, and business-effect disposition. If a reentered state invokes changed code, the actual implementation basis belongs in that result. Preserve contradictions, unexpected visits and unresolved effects. Another redrive is a separate decision, not automatic test cleanup.

Gate: extra effects, wrong identity, unexpected state paths, expired authority, deadline/cost breach or missing telemetry trigger the approved stop and owned investigation. A green workflow result does not clear a failed effect assertion.

12. Stop, investigate and clean up without claiming reversal

Owner: stop operator and application owner. Output: stop/readback and remaining-effect record. Stop further redrive admission first. Use only the separately authorized StopExecution route for the exact Standard execution when the agreed condition requires it. Record request and readback separately. Do not expose credentials, sensitive error content or broaden role privileges to make cancellation easier.

Stopping orchestration is not proof that a request already dispatched was undone. The integration reference documents best-effort cancellation of .sync work and possible continued charges; that construct is excluded here, not silently assumed safe. Even an inert request-response fixture needs a ledger of dispatched/confirmed/unknown operations when stopped.

Any compensating change requires its own business authority and evidence. In this sandbox, cleanup can mean removing separately identified temporary fixture data after the evidence-retention owner approves it. Do not delete execution history, operation receipts or source manifests to hide failed expectations. Do not describe a stopped execution as rollback of its effects.

Gate: unresolved work remains owned with a deadline and next evidence action. Failure to stop or read back becomes an escalation, never a reason for repeated uncontrolled mutation.

13. Hand off the state manifest and a defensible disposition

Acceptance criteria for one redrive decision

The execution owner and destination owner record PASS, FAIL or HOLD against each item and link the actual evidence. An accepted API request is not accepted business settlement.

  • The exact Standard execution, original input/definition and current eligibility are retained. An unsupported Map/integration path or ambiguous history remains HOLD, not a new-execution fallback.
  • Every potentially re-entered visit has a history-derived identity, enclosing error, mutable dependency inventory and destination effect disposition. A state name alone cannot identify a visit.
  • Current business permission, deduplication/guard retention and recovery scope are accepted for each relevant effect. Unknown destination settlement blocks repetition even if the orchestration appears eligible.
  • Retry, loop and enclosing-state exposure is bounded as calls, time and possible effects; fixtures include ambiguous results and failures. The approved cost/time ceiling and stop owner precede any mutation.
  • Any authorized request preserves its identity and uncertain-response reconciliation. The final manifest separates request acceptance, orchestration outcome, destination settlement and retained cleanup obligations.

Close as accepted observed recovery, failed observed recovery, held decision, or unexecuted preparation. Only observed destination evidence can support the first two dispositions; neither a synthetic fixture nor this procedure establishes a recovered customer outcome.

Owner: accountable application owner. Output: accepted tested scope or named gap. Copy this record for one identified execution. The author supplies no populated AWS result. Each actual field requires its referenced evidence; blanks remain UNKNOWN.

Manifest ID / revision / accountable owner / source-check time:
Account / Region / execution ARN / workflow type:
Original start-stop dates / status / error / redrive status-reason:
Definition-revision reference / version-alias / input fingerprint:
Complete history page inventory / prior redrive count-date:
State visit: full path + branch/iteration + history event IDs:
  original disposition / enclosing error / reentry classification:
  original state input reference / Task resource-code/config basis:
  possible effects / operation identity / prior destination evidence:
  matching retries + parent repetition / bounded call exposure:
  current authority-expiry / effect guard lifetime / reviewer:
  expected visits-effects / observed result or UNKNOWN:
Unsupported construct / handoff owner / reconciliation HOLD:
Reader / mutation caller / execution role / stop authority:
Approved fixture scope / token-request identity / deadline / cost ceiling:
Request response-readback / appended events / actual effect disposition:
Stop evidence / dispatched pending work / compensation authority:
Cleanup-retention decision / gaps / owner acceptance / next action:

Done means the authorized fixture has an explained state path and effect disposition, or a transparent held result with an owner. It does not mean every original customer task was recovered. Reconcile the manifest to preserved, observed repeat, newly reached, suppressed, rejected and unresolved work without counting a task in several effect outcomes. Record the actual artifact revision reviewed and all unsupported scope.

The next action is one read-only execution inventory and a visit-level manifest. Where the record exposes missing effect evidence or current authority, commission that reconciliation before considering mutation. A reliability review can help define the evidence and stop boundaries; it is not a replacement for the organization's execution permission. Reading this playbook does not require an email or authorize Ampity to contact the reader.

Related services