Recovering AI-Initiated Business Actions: Identity, Reconciliation and Authority

Design recovery for AI tool calls that change business state. A practical framework for uncertain writes, operation identity, provider contracts, approvals,...

audience="Engineering leaders, application owners and platform teams permitting AI assistants to change records, send messages or initiate business transactions." decision="What the system may do after a write times out, how it proves the actual outcome, and who may authorize a retry or a corrective action." position="Recovery belongs to a durable execution boundary, not to the model's interpretation of an error. Preserve the business operation, reconcile uncertainty and recheck authority before another side effect." scope="An illustrative design for externally visible actions. It is not a claim that every provider supports idempotency, cancellation, authoritative lookup or exactly-once business delivery." outputs={[ 'An operation and attempt identity contract', 'A provider capability matrix', 'An uncertain-write recovery policy', 'An approval and cancellation boundary', 'A failure-injection acceptance matrix', 'A measurable recovery backlog', ]} />

Executive summary

An assistant creates a purchase request. The downstream system commits it, but the connection closes before the assistant receives the response. The model sees a tool error and proposes trying again. A second request uses a new identifier, and the business now has two purchase requests. Nothing in the language model needs to malfunction for this incident to occur. The failure is an execution contract that treats missing evidence as a failed business action.

This paper proposes a recovery boundary for AI-initiated writes. Trusted application code assigns a stable business-operation identity, validates the intended effect and records the dispatch attempt. When evidence is incomplete, the operation becomes unresolved. The executor reconciles it through a documented provider capability or puts it into an accountable review queue. The model may explain the situation or propose a next step, but it cannot manufacture proof that the first write failed.

The recommendation is deliberately narrower than a universal agent platform. Start with one consequential tool, map its real failure behavior and make every recovery transition testable. Add broader automation only after the team can demonstrate how lost responses, repeated deliveries, revoked approvals and partial completion behave. The diagrams show a proposed design, not a customer implementation or a measured Ampity result.

Scope, definitions and assumptions

The design covers actions whose effects outlive a model response: changing a CRM record, publishing a schedule, initiating a payment, sending a customer message or updating access. Read-only queries can still expose sensitive information and consume money, but they have a different recovery problem. This paper focuses on externally visible writes and their subsequent verification.

An operation represents one authorized business intent. An attempt is one effort to execute it. A run is an assistant session that may contain several operations. A receipt is evidence from the authoritative system about a particular effect. Reconciliation compares the intended operation with authoritative state. Compensation is a separate business action intended to address an earlier effect. Unknown means available evidence cannot yet establish the outcome.

Assume that the application can keep a durable operation record and that someone owns unresolved cases. Do not assume distributed transactions across every system. Do not assume a provider exposes reliable search, preserves keys indefinitely or supports undo. If those assumptions fail, automation must become more conservative. The constraints are design inputs, not inconvenient details to hide behind an optimistic retry policy.

Begin with the business consequence

Classify a tool by what it changes, not by whether its API call looks simple. Creating a draft note may be recoverable with modest review. Sending an offer to a customer can create expectations that deleting a database row will not reverse. Changing permissions can expose data before a later correction. A small payload can therefore require stronger controls than a large read-only analysis.

For each action, name the affected party, the resource, the authorized change and the evidence that establishes completion. Then identify the worst credible consequence of duplication, omission, delay and incorrect targeting. A recovery policy should be reviewed by the person responsible for that business consequence. Infrastructure engineers can explain timeouts; they cannot independently decide whether repeating a commercial commitment is acceptable.

Consider an illustrative supplier-onboarding assistant. It may draft a supplier record and recommend a payment contact. It should not infer permission to activate payments merely because a draft was approved. The approval boundary, operation boundary and recovery boundary must describe the same intended effect. Otherwise a technically correct replay can execute a business action the user never authorized.

Separate operation identity from model runs

Assign the operation identifier in trusted code before the write becomes eligible for dispatch. Preserve it across worker restarts and assistant retries. Give each dispatch a separate attempt identifier. Record the model run for explanation and tracing, but do not make the run identifier the sole deduplication identity. A new conversation can refer to the same unfinished operation, while one conversation can legitimately request two different actions.

Store a canonical representation of the approved intent or a secure reference to it. Bind its hash to the operation, tenant, target resource, action type and approval. The canonicalization rules matter: field ordering, omitted defaults and normalization should not accidentally turn equivalent requests into new business intents. Conversely, changing an amount, recipient or effective date must not silently reuse the old operation.

Treat client-supplied keys as scoped input, not as globally trustworthy identifiers. A key reused by another tenant must not reveal that tenant's receipt. A user clicking twice may mean an accidental duplicate, or may mean two deliberate purchases. The interface needs an explicit way to distinguish those cases. Deduplication should preserve intent, not guess it from a convenient text similarity score.

Specify the durable operation record

The ledger is the recovery team's source of continuity. It should preserve enough information to distinguish a missing acknowledgement from a changed business request without becoming an uncontrolled copy of all customer data. Keep sensitive payloads in appropriately protected storage, and retain references or hashes in routine operational views where possible.

| Field | Why it is needed | Review concern | |---|---|---| | Operation ID and tenant | Identify one scoped business intent | Prevent cross-tenant key collisions and receipt disclosure | | Action and target | Define the intended side effect | Avoid broad or ambiguous tool names | | Intent hash and version | Detect altered requests | Document canonicalization and approved defaults | | Approval reference | Bind authority to the exact action | Include expiry, limits and relevant resource version | | Provider and capability version | Select the supported recovery path | Record retention and lookup semantics | | Attempt IDs and dispatch times | Explain repeated network requests | Separate application attempts from SDK retries | | State and revision | Control concurrent transitions | Reject stale workers and invalid state changes | | Receipt reference | Prove the observed effect | Verify target, amount or relevant business fields | | Recovery owner and deadline | Make unresolved work actionable | Avoid an unowned permanent queue |

Persist the intent and authorization decision before dispatch. A local transaction can make local ledger updates atomic, but it does not make the external effect atomic with them. A crash between those two systems remains possible. Design the recovery path for that gap instead of describing a local database commit as an end-to-end guarantee.

Represent uncertainty as a real state

Use explicit operation states that tell the next worker what evidence exists. A useful minimum distinguishes ready, in flight, unknown, verified, definitively rejected and held for review. Cancellation and compensation need their own records or substates. These names are illustrative; the important property is that uncertain completion cannot be confused with a safe-to-repeat rejection.

Record the reason for entering unknown and the evidence required to leave it. A network timeout, an incomplete receipt and a contradictory provider response are different investigations. Give held operations an owner and a next review time. Terminal states should have explicit criteria. A dashboard showing every request as either green or red is simpler, but it conceals exactly the uncertainty that can produce duplicate actions.

Recover the lost-response case

The following sequence assumes a provider can authoritatively locate an effect using a stable operation reference. That is a capability to verify, not a feature to invent for every integration. If search is eventually consistent, an empty result immediately after a timeout is not enough to prove absence. The recovery procedure must account for the provider's documented visibility behavior.

Match the recovered receipt to the original tenant, target and intent. Finding any record with a similar name is insufficient. A receipt with the wrong amount or resource version is a conflict, not success. Persist the verified receipt before telling the reader that the action completed. If verification is delayed, show a truthful pending status and an operation reference that a later session can use.

Choose recovery from the provider contract

Build a capability matrix for each write endpoint. Record whether it supports an idempotency key, how keys are scoped, which payload differences are rejected, how long records remain available, what happens during concurrent requests and how an existing result can be retrieved. Test sandbox behavior where available, while recognizing that a few successful tests do not replace the documented contract.

| Provider capability | Candidate recovery | What stops automation | |---|---|---| | Stable key with documented replay semantics | Repeat the same approved operation within that contract | Key expired, intent changed or authority revoked | | Reliable lookup by unique operation reference | Reconcile first; repeat only after absence is sufficiently established | Visibility lag, ambiguous matches or incomplete fields | | Conditional update with resource version | Read current state and evaluate the original precondition | Version changed or approval no longer describes current state | | No reliable replay or lookup | Hold for accountable review | Missing evidence cannot be repaired by model confidence |

AWS describes caller-provided request identifiers as a way to preserve intent in its idempotent API design discussion. That supports the identity principle; it does not establish a guarantee for an unrelated provider. The matrix above is a proposed engineering review artifact. Populate it with the actual endpoint contract before enabling recovery writes.

Do not confuse HTTP semantics with business safety

RFC 9110, section 9.2.2 defines HTTP idempotency in terms of the intended effect of repeated identical requests. It also cautions against automatically retrying non-idempotent methods without evidence that doing so is safe. These semantics are useful, but an HTTP method alone does not explain every downstream business consequence.

A PUT that sets a record to an approved status may be idempotent at the resource level. If surrounding code sends a new email for every received request, the complete customer experience can still contain duplicates. A POST may support safe repetition through a provider's explicit operation key. Review the whole effect chain, including notifications, webhooks, billing and downstream automation, instead of classifying risk by method name alone.

Use a concrete acceptance statement: repeating this exact operation within this defined window cannot produce a second specified business effect. List the excluded effects and the evidence used to test the statement. Avoid the phrase exactly once unless the boundary, failure model and supporting mechanism are explicit. A broad slogan creates confidence precisely where the review needs careful qualification.

Account for key expiry and cached errors

Stripe's idempotent request documentation provides a useful example of a specific contract: it describes retained results, parameter checks, error behavior and key pruning. The lesson is to read those details for the actual endpoint, not to assume every payment, CRM or scheduling API behaves the same way.

A retry after a deduplication window expires can become a new effect even if the application preserved the original key. Store the window and the last safe recovery deadline alongside the operation. Do not use an arbitrary business deadline as a substitute. If a job resumes after a prolonged outage, it must evaluate the remaining provider guarantee before submitting anything.

Some providers can return a stored error for a repeated key. Changing the key just to obtain a different response may bypass duplicate prevention while leaving the original outcome unresolved. Treat a cached error according to the documented contract and reconciliation evidence. If there is no supported way to establish the effect, escalate it as unresolved. An apparent retry workaround is not a recovery design.

Control concurrency and stale workers

Two workers can receive the same queue message, or one can resume after its lease expired. A process-level mutex does not protect separate hosts. Use a durable claim or conditional state transition, and attach a revision or fencing value where the execution boundary can enforce it. The worker that no longer owns the operation must not overwrite the newer worker's decision.

Local claims alone do not prevent an already-dispatched external request from completing. If worker A loses its lease while the provider is processing its request, worker B should reconcile the operation rather than treating the expired lease as proof of failure. Provider-side operation identity remains important. Where the provider cannot honor a fencing value, document that limitation explicitly.

Test the moment between claim acquisition and dispatch, the moment after dispatch and the moment before receipt persistence. Include delayed responses from an old worker arriving after a newer reconciliation result. An immutable attempt history helps investigators understand the sequence. Updates to the current state should still use guarded transitions, so late evidence cannot silently downgrade a verified result or reopen a cancelled business intent.

Bind recovery to current authorization

Idempotency is not permission. An action can be safe from duplication and still be unauthorized at the time of a retry. Approval should bind the actor, tenant, action, target, relevant values, resource preconditions and expiry. Store that binding outside the model context. A narrative saying the user approved it is not enough to authorize a write.

Before a recovery write, recheck whether the original approval remains valid and whether the requested effect is unchanged. If a user approved a price update for a particular product version, a subsequent catalog change may invalidate the precondition. The recovery service should not reinterpret the approval to cover the new state. Create a revised proposal and ask for the appropriate authorization when required.

Separate read authority for reconciliation from write authority for repetition or compensation. A system may be allowed to check whether a payment exists without being allowed to issue another payment. The model can present the uncertainty and request a decision, but the executor must independently enforce the result. This separation also makes permissions easier to test during incident response.

Treat cancellation as a request, not an undo

The MCP cancellation specification, version 2025-06-18 describes an optional cancellation notification with timing and handling limitations. A cancelled tool interaction therefore cannot, on its own, prove that a remote business effect never happened. This paper cites that specific version rather than claiming a universal behavior across current implementations.

Provide different user-facing outcomes for cancellation before dispatch, cancellation requested during execution and cancellation confirmed by the authoritative system. After dispatch, keep the operation discoverable until the outcome is reconciled. A cancelled browser request or abandoned chat window should not erase the ledger record. Otherwise the business loses the evidence needed to address an effect that completes later.

If an action is verified after the user asked to cancel, apply the business policy for that situation. It may require a compensating action, a notification or a review. Do not let the model improvise the response based on a friendly conversational promise. The interface should say what was stopped, what remains uncertain and when the user can expect a verified update.

Make compensation a separately authorized operation

Compensation does not necessarily restore the world to its earlier state. Deleting a duplicate record cannot make a recipient unread an email. Cancelling a booking may incur a fee or release capacity that someone else immediately takes. Refunding money is another transaction with its own failure behavior. Define what correction is possible and which residual consequences remain.

Microsoft's compensating transaction pattern emphasizes application-specific compensation and the possibility that compensation itself fails. Use it as a pattern reference, not a promise of automatic reversal. The corrective operation should have its own identity, authority, evidence and recovery procedure, linked to the original effect.

For the supplier-onboarding example, disabling an accidentally activated supplier might be appropriate only after checking pending transactions and ownership. Removing it without that check could worsen the incident. Record the rationale for the selected correction and any manual steps. Close the original recovery case only when the agreed outcome and remaining risks have been acknowledged by the responsible owner.

Plan partial completion across several services

One assistant request can contain several effects: create a record, attach a document, schedule a task and send confirmation. Avoid representing the entire chain with one undifferentiated success flag. Give each effect a step identity, receipt and recovery decision. The overall operation can then be partially complete without forcing the executor to repeat the successful steps.

Choose an ordering based on consequence. A notification should normally reflect durable business state, not an optimistic intermediate result. If a required attachment failed, decide whether the record remains a draft, whether later repair is acceptable and whether the user needs to know. These are workflow decisions that should be agreed before release, rather than discovered through improvised model explanations.

An outbox or durable workflow can coordinate local state and planned messages, but it does not remove every external uncertainty. Review delivery semantics at each boundary. A consumer receiving the same event twice still needs a safe effect contract. Distinguish successfully handing a message to a queue from the downstream business action completing. The receipt model should make those stages visible without overstating what any one component guarantees.

Set retry budgets across all layers

Count retries in the assistant loop, workflow engine, queue, HTTP client and SDK. If each layer repeats independently, the number of downstream attempts can exceed what a team expected from a single retry setting. Assign one owner to decide when another business dispatch is allowed, while still accommodating documented transport-level behavior inside a client.

AWS SDK retry behavior documentation explains configurable attempts and retry strategies. Its current page also distinguishes new 2026 behavior from earlier settings. Check the installed SDK and configuration rather than copying a default from a different version. This paper intentionally specifies no universal delay, attempt count or environment setting.

For each operation, define a dispatch budget, reconciliation budget, elapsed-time limit and escalation threshold. Exhausting the budget should lead to a meaningful pending or held state, not a new operation identifier. Backoff reduces pressure on a struggling dependency; it does not prove a write is safe. Recovery policies need both a load-control decision and an effect-safety decision.

Make security and operational evidence work together

Recovery records can expose customer information, commercial values and privileged action details. Restrict access by tenant and role, redact routine logs and protect the underlying evidence store. A support view should not show secrets merely because they appeared in a tool payload. Record access to sensitive evidence where necessary, and define retention separately from convenient debug logging.

Capture operation ID, attempt ID, policy version, provider reference, state transition and a reason code in structured events. Preserve enough timing to distinguish dispatch, provider acknowledgement and verification. Avoid relying exclusively on a model transcript: it can be verbose, omit important transport facts and contain untrusted content. Operational decisions should be reconstructable from trusted execution evidence.

Treat provider responses and retrieved records as untrusted data for any subsequent model interaction. A record containing instructions must not acquire authority to trigger another write. The recovery service should validate typed fields and permitted transitions before exposing a summarized result. Security review should include forged receipts, cross-tenant references and attempts to use a recovered record as authorization for a new action.

Test the uncertainty, not just the happy path

Create a controlled test provider or sandbox fixture that can commit an effect while withholding its response. This is the essential case that ordinary mocked errors often miss. A mock that throws before doing anything only proves how the application handles a known non-effect. It does not demonstrate protection against uncertain completion.

| Injected condition | Required evidence | Release fails if | |---|---|---| | Provider commits, response is dropped | Unknown state followed by matching receipt or safe hold | A fresh operation or duplicate effect is created | | Worker crashes after dispatch | Restart recovers the original operation | Restart assumes absence and repeats blindly | | Two workers receive the same task | One operation, guarded transitions and explainable attempts | Both independently authorize a new business action | | Lookup temporarily returns no result | Policy respects visibility uncertainty | Empty search is treated as definitive absence | | Approval expires during recovery | Read reconciliation may continue; new write is prevented | Old approval permits changed or expired action | | Key retention window has elapsed | Held state or supported alternative evidence | A presumed replay becomes a new write | | Compensation times out | Corrective operation remains independently recoverable | Original incident is falsely reported as resolved | | Old response arrives late | Receipt is correlated and transition rules hold | Stale worker overwrites verified state |

Review actual provider requests and durable records after each test. A green application assertion alone may miss a duplicate external effect. Use synthetic tenants and non-delivering message endpoints where possible. Document the limits of the test environment and reserve destructive production exercises for explicitly approved procedures. A recovery test should not become an accidental customer communication.

Measure recovery as a business process

Track unresolved operations by age, action type, consequence and owner. Measure time to verified outcome, duplicate effects, unverified success messages, review workload and compensation failures. Keep those measurements separate from tool-response success rates. A system can return fewer errors while silently repeating more business actions, which would be an unacceptable trade.

Estimate cost using the complete recovery path: model steps, provider requests, lookup traffic, evidence storage and human review. Compare it with the cost of a duplicate or omitted action. The right policy for a low-consequence draft may differ from one for a customer-visible commitment. Do not force every tool into the same budget merely because that simplifies a central configuration screen.

Define targets using an observed baseline and the business deadline. An arbitrary five-minute objective is not automatically suitable for every action. Some operations require immediate blocking and review; others tolerate asynchronous reconciliation. Alert on ownership gaps and backlog growth as well as absolute age. A durable queue without a working consumer and an accountable reviewer is stored failure, not completed recovery.

Introduce the boundary without rewriting the product

Start by inventorying the few tools that currently change external state. Select one whose duplication or omission would be visible to customers, and write its capability matrix. Add operation IDs and evidence capture before enabling automatic corrective behavior. This makes existing ambiguity observable and gives the team real cases to use in tests.

Next, route uncertain outcomes into a manual review process with a clear owner. Measure whether reviewers can find the relevant provider state and make a defensible decision. Only then automate a narrow transition, such as reconciling an existing receipt. Keep the dispatch authority restricted while verifying the recovery path under worker crashes and repeated deliveries.

Expand action by action. Each additional provider must earn its recovery policy through documentation and evidence. Reuse the ledger and review interface, not assumptions about provider semantics. Roll out behind a capability switch so a failing recovery rule can be disabled without losing the underlying operations. Preserve visibility into pending work during rollback, and define who reviews cases created under an older policy version.

Limitations and cases requiring a different choice

This design does not guarantee exactly-once outcomes across arbitrary third-party systems. Some providers offer no authoritative lookup, cannot correlate a stable reference or retain deduplication state for too short a period. If the consequence is unacceptable and the evidence is insufficient, the correct choice can be draft-only assistance or human execution rather than automatic writes.

An operation ledger adds availability, privacy and maintenance responsibilities. It can itself become unavailable or contain incorrect bindings. Test backup, access, migration and state-transition behavior. For a purely read-only assistant, a smaller execution trace may be sufficient. Do not add a complex recovery platform to every summarization feature without a concrete side-effect problem to solve.

Business judgment can remain necessary even with excellent evidence. A customer may request a new action that resembles an earlier one, or a corrective step may create a new commercial obligation. The model can help assemble relevant facts, but it cannot replace the authorized owner of that decision. Automation should reduce unexplained uncertainty, not hide difficult judgments behind a confidence score.

Decision checklist and next step

Before authorizing recovery writes, review one completed example with the application owner, integration engineer and person accountable for the business effect. Can they identify the exact intent, every dispatch, the valid approval and the final receipt without reconstructing events from memory? If not, improve that evidence before increasing autonomy.

  • Define one operation independently from runs and attempts.
  • Bind it to tenant, target, intent version and approval.
  • Persist the record before dispatch and guard state transitions.
  • Verify provider replay, lookup, concurrency and retention semantics.
  • Keep unknown distinct from rejected and safe to repeat.
  • Recheck authority before a retry or compensation.
  • Give unresolved cases an owner, deadline and review procedure.
  • Test lost responses after commitment and inspect external effects.
  • Show readers verified, pending and held outcomes truthfully.
  • Review recovery cost and backlog alongside business consequences.

The next useful artifact is a filled provider capability matrix and one lost-response test, not another generic agent diagram. Bring those results to an agentic workflow review if you need help deciding which transitions are safe to automate. For the broader boundary around tools, retrieval, evaluation and release, see AI agent architecture in production.

Primary references and reading boundaries

The companion write-recovery audit playbook supplies the owner-led procedure, isolated test setup and acceptance evidence for this design. For a shorter introduction to the lost-response problem, read AI tool timeouts and duplicate business actions.

The linked sources support specific protocol and provider statements where they appear. The ledger schema, diagrams, review matrix and supplier scenario are proposed designs in this paper, not extracts from those sources. Endpoint behavior, SDK defaults and protocol versions should be rechecked during implementation. No source establishes that an unrelated CRM, payment or scheduling integration has the same guarantees.

Use the paper as a working review framework. Replace each illustrative assumption with your application's actual contract and record what remains unknown. A good architecture review can legitimately end with a narrower automation boundary. That is a useful result when it prevents the system from repeating an action it cannot safely explain.