Audit an AI Tool’s Write-Recovery Path
Run a controlled recovery audit for an AI tool that changes business state. Produce an operation inventory, lost-response evidence, concurrency tests, approval checks...
trigger="An assistant can create, update, publish or send something, but the team cannot explain what happens when a write times out or a worker restarts." owner="The application owner accountable for the business effect." participants={['Integration engineer', 'Platform engineer', 'Security reviewer', 'Business process owner']} prerequisites={[ 'One named write tool and its actual provider contract', 'A synthetic tenant and isolated, non-delivering test endpoints', 'Access to application attempts and independent provider-effect records', 'Permission to inject failures in the test environment', ]} outputs={[ 'A scoped operation and provider capability record', 'A reproducible lost-response fixture and evidence packet', 'Concurrency, approval and cancellation test results', 'A reconciliation runbook with an owner and stop conditions', ]} doneWhen={[ 'Each required fault has evidence of actual effects, not just tool responses', 'Unknown outcomes cannot silently become fresh writes', 'Expired authority blocks recovery writes', 'Unresolved cases have a reachable owner and a tested review path', ]} />
Use this when a tool can leave a lasting effect
A retry audit should answer a concrete question: can this approved operation be recovered without creating a second unintended business effect? It is not an agent accuracy benchmark. The model can choose the correct tool and still produce a duplicate when a downstream commit succeeds but its response is lost. Start with the execution boundary that changes external state, not with a new prompt telling the model to be careful.
Choose one consequential tool for the first audit. A CRM create, schedule publication or customer notification is a reasonable candidate when its outcome matters to the business. Do not combine every write-enabled tool into one generic pass result. Provider contracts, approval rules and evidence sources differ, so each tool needs its own acceptance record.
The test harness and examples below are proposed procedures using synthetic data. They are not claims about an Ampity customer deployment. Never run message-delivery, payment or permission tests against real customers merely because a test account exists. Confirm which actions can escape the sandbox before starting.
1. Name the effect and appoint its owner
Owner: application owner, with the business process owner. Output: one operation contract. Describe the intended effect in a sentence that a non-engineer can verify. For example: create one draft supplier request for this synthetic tenant and reference. Specify whether a notification, task or downstream webhook is part of the same accepted outcome or a separate step.
Write the duplication, omission and delay consequences. Include effects outside the primary record store. Deleting a duplicate request may not retract an email generated when it was created. This distinction determines which observations and corrective procedures the audit needs.
Set a scope boundary around the first test. Record the application version, tool endpoint, provider environment and relevant feature configuration. Identify the person who can disable the write capability if the audit fails. If nobody can make that decision, stop and resolve ownership before injecting faults.
2. Collect the real provider capability contract
Owner: integration engineer. Output: a completed capability matrix with source links. Read the documentation for the exact write endpoint and SDK version. Record key scope, parameter matching, retention, concurrent request behavior, status lookup and eventual visibility. Mark unknown fields as unknown rather than filling them from experience with another provider.
| Capability | Recorded answer | Required evidence | |---|---|---| | Stable request key | Supported, unsupported or unknown | Exact endpoint documentation | | Key retention | Documented interval and exclusions | Provider statement and tested boundary where feasible | | Existing result lookup | Authoritative reference or search only | Fields and visibility behavior | | Concurrent duplicates | Reject, serialize, replay or unknown | Contract plus isolated test | | Changed payload with same key | Rejected, accepted or unknown | Parameter-binding contract | | Cancellation | Before dispatch, in progress or unsupported | Actual implementation behavior |
The HTTP idempotency definition helps classify repeated requests, but it does not establish the complete business workflow's safety. Stripe's documented idempotent requests illustrate why stored errors, parameter checks and key pruning belong in the matrix. Those details are provider-specific. Do not copy Stripe's retention assumptions into an unrelated CRM test.
3. Trace operation identity through every layer
Owner: integration engineer, with the platform engineer. Output: an identity trace. Follow one test intent from the interface through the assistant, tool adapter, worker, client and provider. Record the operation identifier, each attempt identifier and the provider reference. Confirm which component creates each identifier and whether it survives a restarted process.
Check whether a model retry, queue redelivery or SDK retry changes the key. Distinguish a new legitimate business request from a second attempt at the old one. Payload similarity alone cannot settle that distinction. AWS's idempotent API discussion explains the value of expressing caller intent through request identifiers; it does not give your integration a guarantee without a compatible endpoint contract.
Capture a redacted sample trace that joins the layers. Do not use an email address or sensitive account value as the key. Scope identifiers to the tenant and action. The pass condition is that an investigator can connect all attempts to one intent without reading the entire model conversation or making a guess based on timestamps.
4. Prepare independent observations
Owner: platform engineer. Output: a test harness and reset procedure. The application record tells you what the executor believes happened. An independent provider-effect record tells you what actually committed. You need both. A stub that returns an error before creating anything does not test uncertain completion.
Configure the controlled provider to keep a durable or inspectable commit log even when it drops a response. Ensure the executor uses its real recovery code, rather than a simplified path that bypasses workers, persistence or authorization. Record the differences between the harness and the real provider. An isolated harness proves your application's response to modeled faults; it cannot prove undocumented third-party behavior.
5. Establish the normal-operation baseline
Owner: test engineer or integration engineer. Output: a complete baseline evidence packet. Submit one approved synthetic operation with no injected failure. Observe the intent, dispatch, provider commit, receipt verification and reader-visible status. Confirm the resulting record's target and fields, not only that some record exists.
Count the downstream business effects included in the operation contract. If the tool creates a record and a message, inspect both. A duplicate message should fail the baseline even when the record count is correct. Keep network request count separate from business-effect count; legitimate transport retries may produce several requests for one effect.
Reset the harness using its documented test procedure and repeat with a different legitimate operation identifier. Confirm that the system permits two deliberately distinct intents. A deduplication rule that suppresses real work is not safer simply because it produces fewer effects. Preserve both baseline packets for comparison with failure cases.
6. Inject a lost response after commitment
Owner: platform engineer executes; integration engineer inspects. Output: the core uncertain-write test. Configure the provider fixture to commit the effect and then withhold or interrupt the response. Use a barrier or explicit fixture signal so the timing is deterministic. A sleep that occasionally lands after commitment produces a flaky demonstration, not a reproducible acceptance test.
Observe the executor's first state after the connection fails. It should preserve the original operation and represent the missing outcome evidence. Then allow its documented reconciliation or safe replay path to run. Inspect the independent effect log after recovery. Require the expected effect count, correct target and a receipt matching the original intent.
The following is test pseudocode, not a runnable provider integration:
fixture = provider.commit_then_drop_response(operation_reference)
executor.submit(approved_intent, operation_reference)
wait_until(fixture.commit_observed)
wait_until(executor.records_uncertain_outcome)
executor.recover(operation_reference)
assert provider.effects_for(operation_reference) == expected_effects
assert executor.receipt_matches(approved_intent)
assert executor.did_not_create_a_fresh_operation()Define bounded waits and diagnostic output in the actual harness. A timeout waiting for evidence is an unresolved test, not a pass. If recovery cannot safely continue, a held operation with an owner can be the correct result. Do not force every provider into an automatic-success expectation.
7. Restart the worker at the commit boundary
Owner: platform engineer. Output: a restart recovery trace. Stop the test worker after dispatch but before it persists a verified receipt. Use the fixture's commit signal to select the boundary. Start a replacement worker using the same durable operation storage and queue configuration that the deployment uses.
Inspect whether the replacement discovers the existing operation, preserves its key and reconciles its effect. A fresh run identifier is acceptable if it remains attached to the original business operation. A fresh business key introduced because the process forgot the old one is not acceptable for an uncertain write.
Also stop the worker before dispatch. This separate case should not create a provider effect. Together, these tests demonstrate that the system distinguishes a known unsubmitted operation from an uncertain submitted one. Keep evidence from both boundaries; a successful restart before dispatch does not prove recovery after commitment.
8. Exercise concurrency and delayed evidence
Owner: integration engineer. Output: concurrency and stale-worker results. Deliver the same operation to two workers under controlled scheduling. Let both reach the claim boundary, then release them in a reproducible order. Inspect durable claim revisions, dispatch attempts and provider-effect records. Do not conclude safety from a mutex that protects only one process.
Simulate lease expiry while an earlier request is still running. The replacement worker must not interpret the expired lease as proof that no effect occurred. Then release the earlier worker's delayed response after a later reconciliation result. Verify that guarded transitions prevent stale evidence from overwriting a newer verified state.
Record the contract limitation if the provider cannot enforce fencing or concurrent-key semantics. The acceptable response may be a conservative hold rather than a second dispatch. The purpose is to identify that boundary before release, not to make every test green by adding progressively broader retry rules.
9. Change authority while recovery is pending
Owner: security reviewer, with the application owner. Output: authorization evidence. Create an approval for a specific synthetic target and intent. Inject uncertain completion, then expire or revoke the approval before another write attempt. Permit scoped read-only reconciliation if the application policy allows it, but verify that a new write cannot use the expired authority.
Repeat after changing a relevant resource version or approved value. An approval for one amount or recipient should not silently authorize a different one. Inspect enforcement in trusted execution code rather than relying on a prompt telling the model to request permission again.
Record the denial reason and reader-visible outcome. A generic tool error may cause the assistant to seek a different path to the same unauthorized effect. The denied state should survive worker restarts and new conversations. Review alternative tools that could bypass the intended boundary, and disable those paths until they enforce the same authorization contract.
10. Verify cancellation without promising reversal
Owner: integration engineer. Output: cancellation outcome matrix. Test cancellation before dispatch, during provider execution and after commitment. Record whether work stopped, whether an effect exists and whether the user-facing status accurately describes the result. The MCP cancellation specification dated 2025-06-18 describes optional cancellation and timing races. Do not treat a sent notification as evidence of a business undo.
The executor must retain a discoverable operation after the chat or browser is closed. If the provider completes later, the system still needs to reconcile that effect and apply the agreed business policy. Test the account or tenant that will perform that reconciliation, including its read authority and access to the evidence.
If correction is required, model it as a separate authorized operation linked to the original. A test passes only when the correction's outcome is verified or explicitly unresolved with ownership. Deleting the local operation record to make the interface look cancelled should fail the audit.
11. Test the limits of lookup and retention
Owner: integration engineer. Output: recovery-window and ambiguity results. Configure the harness to delay lookup visibility after commitment. Verify that an empty search does not immediately become proof of absence. Then return a similarly named record with the wrong operation reference, tenant or approved fields. The executor should reject it as insufficient evidence.
Exercise the documented replay window. Where waiting for real expiry is impractical, use a controlled clock or provider fixture and label the evidence as modeled. Do not describe that modeled test as a measurement of the production provider. Preserve the last safe replay deadline in the operation record.
Test what happens when the key is no longer protected by the provider contract. A held case or supported alternative reconciliation can be acceptable. A retry with a new key merely to escape a stored error is not. Also check changed parameters with the same key, so the system cannot confuse a new business intent with recovery of the old one.
12. Write the reconciliation runbook
Owner: application owner. Output: a usable reviewer procedure. Give reviewers a stable operation reference, permitted evidence sources, scoped credentials and a decision table. They should know which facts establish success, which establish non-execution and which leave uncertainty. A queue entry saying only that an error occurred is insufficient.
| Observed evidence | Permitted next action | Decision owner | |---|---|---| | Matching authoritative receipt | Record verified outcome | Authorized recovery service or reviewer | | Proven non-execution, current approval and safe contract | Repeat the original operation | Executor under the approved policy | | Ambiguous lookup or expired replay window | Hold and investigate | Named application owner | | Wrong target or conflicting effect | Open corrective-action review | Business process owner | | Approval revoked or intent changed | No new recovery write | Authorized approver for any revised action |
Rehearse one held case with the actual reviewer role. Confirm the person can find the evidence, record a decision and identify remaining risks. Set an escalation deadline based on the business consequence. A durable queue with nobody assigned to drain it is not a completed recovery path.
13. Apply the release acceptance criteria
Owner: application owner approves the audit record. Output: a scoped release decision. Use the following checklist with links to test evidence, not just checked boxes. Separate failures in the application from provider capabilities that remain unknown. Either can justify narrowing the enabled tool authority.
- The operation contract names every accepted business effect.
- The endpoint capability matrix has primary-source references and explicit unknowns.
- Operation identity survives retries, restarts and repeated queue delivery.
- A post-commit lost response produces one intended effect or an accountable hold.
- Provider state, not tool status alone, supports verification.
- Stale workers cannot silently replace newer decisions.
- Revoked or altered approval prevents a recovery write.
- Cancellation wording matches the actual evidence.
- Replay-window expiry and lookup ambiguity have tested stop conditions.
- A reviewer can resolve or escalate a held case using the runbook.
Archive the fixture configuration, application version, redacted traces and effect records together. Record the limitations of the evidence. Approval for one tool and provider version is not approval for every integration in the assistant. Repeat relevant cases when changing the adapter, persistence, retry configuration, approval policy or provider contract.
Stop conditions, rollback and useful next steps
Stop the exercise if test effects can reach real customers, if tenant isolation fails, if credentials are broader than agreed or if the team cannot reliably reset the fixture. Disable automatic writes for the affected tool when duplicate or unauthorized effects are observed. Keep read-only assistance or draft generation only where that narrower mode is explicitly safe.
Rollback must preserve unresolved operation records. Reverting code while deleting the recovery queue can leave the business with effects nobody will reconcile. Appoint an owner for operations created under the previous policy and decide whether the replacement code can safely interpret them. Do not bulk-replay old records merely because the deployment was rolled back.
The next useful step is to take one failed fixture and change the narrowest execution boundary that owns it. Rerun that case, the normal baseline and adjacent failure cases before widening authority. For the design behind these checks, use Recovering AI-initiated business actions. For the wider system, use AI agent architecture in production. If your team needs help establishing the operation contract and acceptance evidence, bring the completed audit packet to an agentic workflow review.