What a Write-Enabled MCP Tool Must Say About Retries

Review operation identity, parameter binding, retention and downstream effects before allowing an MCP write tool to retry after an uncertain result.

Define retry behavior before enabling writes

A write-enabled MCP tool needs an application contract for repeating an operation. Specify which identity survives retries, which inputs it covers, how long duplicate protection lasts and how to discover an uncertain outcome. Without those answers, a timeout should hold the task for reconciliation rather than invite the model to try another write.

The contract must describe the business effect. Creating a database record, sending a message and publishing a schedule have different completion evidence. A tool that performs several of those actions needs separate receipts or a clearly defined aggregate outcome. The caller should be able to distinguish an accepted request from a completed effect.

This article proposes a review method for tool owners and platform engineers. The scheduling example and sample fields are illustrative application design, not a mandatory MCP schema or an Ampity customer incident. Review the actual downstream API and deployed server before choosing a retry policy.

Start by naming one tool, one consequential effect and the system that can confirm it. If that system cannot distinguish this operation from an unrelated change, record the limitation before allowing automatic retries. A natural-language success message cannot fill a missing recovery boundary.

Understand the limits of the idempotency hint

The MCP schema reference, version 2025-11-25 includes idempotentHint, which describes repeated calls with identical arguments as having no additional effect. Its annotations are hints, and the specification warns against using annotations from untrusted servers to make tool-use decisions. A hint needs implementation evidence behind it.

A custom operation ID can appear in a tool's input schema, but merely adding the field does not implement duplicate protection. The executor must validate and retain it according to the application contract. Do not infer that MCP automatically gives every tool a durable deduplication store or makes downstream APIs idempotent.

The MCP tools specification for the same version separates protocol errors from tool-execution errors. Neither category proves that an external write never happened. An executor can lose a response after the receiving system committed its change; the caller still needs effect-specific recovery.

Keep the claim consistent with the whole handler. A function that sets a record to the same value but sends another notification on every invocation has an additional effect. Describe that behavior rather than advertising unconditional idempotency. If repetition is safe only for a particular target state or within a retention window, document and test those conditions.

Give one approved operation an identity that survives attempts

Use a business-operation identifier that the application establishes before execution. Retrying the same approved action keeps that identifier. Individual network calls can have different attempt identifiers for diagnostics, but a gateway retry, a worker restart and an agent retry must not silently become three new business operations.

Bind the operation identity to a defined scope: authenticated requester or tenant, action type, target, relevant revision and canonicalized inputs. State which fields affect intent and how defaults are resolved. Canonicalization should reflect actual semantics. Treating a changed recipient or amount as equivalent because a loose hash ignored it defeats the protection.

Do not derive identity solely from similar-looking arguments. Two legitimate operations can have identical inputs, such as two separately approved purchases. Conversely, one operation can pass through several transport envelopes. The AWS Builders' Library discussion of idempotent APIs explains caller-provided request identifiers, parameter-mismatch handling and the need to coordinate token recording with mutation. These are useful design considerations, not a universal guarantee for an MCP wrapper.

Keep authorization separate from duplicate recognition. A retry must not expose another tenant's stored result, and a remembered identifier must not grant authority for a new effect. Define what happens when approval expires while an operation is in flight. Returning an already completed result and starting new work require different policy decisions.

Make mismatches, concurrency and expiry explicit

When an operation ID is reused with different consequential inputs, stop and report a conflict. The caller should not be allowed to replace the stored intent simply because the earlier response was inconvenient. Correcting a proposal requires a new approved operation, with an explicit relationship to the superseded one where the workflow needs it.

Concurrent requests need an execution rule as well. Two handlers can both observe that an identifier is absent and then both write. A check followed by an uncoordinated create is insufficient. Use the receiving system's supported idempotency mechanism or a storage/execution design that serializes the operation and survives the relevant crash windows. A process-local lock does not establish protection across all workers.

State how long duplicate protection remains effective and what survives after cleanup. Retention is part of the contract, not housekeeping hidden from callers. A delayed job may outlive the window. Before repeating it, recovery needs a trustworthy receipt or readback route; issuing a fresh identity can create a second effect.

Stripe's idempotent-request documentation gives a concrete provider-specific example: repeated keys return stored responses, input changes are checked, and reused keys can represent new requests after pruning. Its rules must not be copied as assumptions about other APIs. Read the retention and execution semantics of each destination you use.

Decide how an in-progress response differs from an uncertain outcome. In-progress means a known operation can still be observed. Unknown means the available evidence cannot establish what happened. A failed readback request must not be interpreted as “nothing exists.” Keep the task held if the status service is unavailable, access is denied or matching records are ambiguous.

Worked example: publishing a schedule and notifying coaches

Consider a hypothetical sports platform with a publish_schedule tool. An operator approves schedule revision 42. Publication updates the schedule store and then enqueues notifications to coaches. The application gives that approved operation identity op-42-a; this label is a sample, not a recommended identifier format.

The schedule store commits revision 42, but the handler times out before recording its local completion receipt. The agent sees an error and proposes another call. If the wrapper creates op-42-b for that retry, it has lost the relationship to the original intent. Even if the schedule update is harmless when repeated, notification delivery may run twice.

Recovery first checks the authoritative publication record for the original operation and revision. If publication is confirmed, it records that receipt and examines the notification stage independently. It should not republish the schedule merely because notifications are pending. Conversely, a queue receipt establishes that notification work was accepted, not that every coach received it.

| Observed evidence | Application decision | Evidence still needed | | --- | --- | --- | | Rejection before execution, with no effect possible under the tested handler | Correct the proposal through the approved path | New input and authority checks | | Matching publication receipt, notifications pending | Preserve publication; recover notification work separately | Per-stage delivery state | | Same operation ID, changed revision | Reject conflicting reuse | Approval for the changed proposal | | Timeout and unavailable readback | Hold the operation as unknown | Receiving-system reconciliation | | Old identifier beyond guaranteed retention | Do not blindly replay | Durable receipt or operator-led recovery |

Suppose the source schedule has moved to revision 43 while recovery is underway. Recovery must distinguish confirming the earlier publication from requesting a new publication. The original operation ID belongs to revision 42. It should not be used to publish revision 43 under an approval the operator never reviewed.

Notification recovery also needs audience binding. If the recipient list changed after publication, do not silently expand the original delivery job. Define whether the job uses an immutable audience snapshot or an explicitly versioned recomputation policy. That choice affects who receives a message and what a repeated delivery means.

Return states that the caller can act on

Define application-level result states that fit the effect, then map them consistently to tool output. For example, accepted, completed, rejected, conflict and unknown can be useful distinctions. These names are proposed application vocabulary, not MCP-required fields. Include a scoped operation reference and the next permitted action.

For completed work, specify what was confirmed and where. If a publication is complete while notifications remain queued, return both stage states. Avoid an aggregate success flag that causes the model to tell the user every stage finished. Keep the user-facing explanation aligned with the receipt, not with the optimism of the request path.

For unknown work, make the restriction explicit: query the status mechanism or refer the task to the named reconciliation owner; do not create a new operation. Status access must use trusted identity and object scope. An operation identifier should never become a public bearer token for viewing someone else's results.

Keep stored results bounded. Recovery needs operation identity, intent binding, timestamps, stage receipts and state transitions. It rarely needs complete sensitive source documents in general diagnostic logs. If a result contains protected content, enforce access when returning a remembered result as well as during the initial call.

Test the contract at the receiving system

Use isolated synthetic records and an approved test destination. Send the same operation twice, change a consequential parameter while keeping its identity, and send two concurrent calls. Then inspect the receiving system for the exact number of effects. A caller receiving the same JSON twice is insufficient evidence if two notifications or records were created behind it.

Inject failure after the downstream commit but before the wrapper saves its receipt. Restart the worker, repeat the attempt with the original identity and verify the recovery decision. Test the opposite window too: an operation reservation exists, but no write started. A stuck reservation should have a defined owner and recovery path rather than remain in progress forever.

Exercise retention boundaries and delayed requests. Repeat immediately, near the stated expiry and after protection has ended. Use a controlled clock or dedicated test policy rather than waiting on a production retention setting. Confirm that late requests follow the documented reconciliation policy and do not recreate a deleted or superseded business effect.

Include layered retries in the test. A downstream SDK, proxy and agent can all retry independently. Record attempt counts, operation identities and the time budget across layers. The acceptance record should show how repeated calls converge on the intended effect or stop in an honest unresolved state.

Limitations and the next tool review

An idempotency contract covers a defined operation within defined conditions. It does not guarantee eventual success, prevent every unrelated update or make an entire multi-system workflow atomic. Irreversible external effects and destinations without suitable identities can require operator-led reconciliation. State those limits before enabling autonomous writes.

Start the next review with one page containing the tool version, effect, operation-ID owner, input binding, concurrent-call behavior, retention policy, authoritative readback and unknown-outcome owner. Attach the failure-injection evidence and rejected-mismatch fixture. If any of those fields is unknown, keep automatic write retries disabled for that effect until the gap has a tested resolution.

For the authority boundary, read why tool descriptions do not establish permission. For the interruption path, read tool timeouts and duplicate business actions and use the action-recovery playbook. Explore agentic workflow engineering if your integration needs an enforced operation contract. You can read these resources without submitting contact details.