Thrive AI Engineering
← All questions

agentic-ai · Editorial starter

How do I retry an agent tool call without performing the action twice?

By Thrive AI Editorial · Created 2026-10-07 · Updated 2026-10-07

Thrive AI Editorial · AI-assisted worked example, not a member conversation.

An agent calls a tool to create a work item. The server commits the item, but the client times out before receiving its response. Retrying the call can create a duplicate. How should the application distinguish a retry of the same intended action from a genuinely new action with identical arguments?

This is an editorial scenario, not a report of a member's production incident. The example below isolates the commit/response boundary without a model, network service or paid API.

0 score

Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.

1 published answer

A selected solution is chosen by the question author or moderator, not an independent certification. Check assumptions before using any code in production.

Thrive AI Editorial

Thrive AI Editorial · AI-assisted worked example, not a member conversation.

2026-10-07

Keep the intent stable across transport retries

A timeout tells you that you did not receive a response. It does not prove that the action failed. Retrying the model prompt is also not a reliable way to recover the original tool execution: it may produce a different call or identifier.

Create an operation identifier in your application when the user approves an action. Persist that identifier with the pending work. Use it for every transport retry of that action. A new intentional action gets a new identifier, even if its arguments are identical.

Code / output
approved intent -> stable operation key -> tool service
                                       -> commit result + key atomically
lost response   -> same operation key   -> return stored result

Runnable boundary simulation

Python 3, standard library only. This deliberately single-process example demonstrates the contract, not a production database implementation.

Code / output
import json

class WorkItems:
    def __init__(self):
        self.items = []
        self.receipts = {}

    def create(self, caller, key, title):
        scope = (caller, key)
        fingerprint = json.dumps({"title": title}, sort_keys=True)
        if scope in self.receipts:
            previous, item_id = self.receipts[scope]
            if previous != fingerprint:
                raise ValueError("same key, different intent")
            return item_id
        item_id = len(self.items) + 1
        self.items.append({"id": item_id, "title": title})
        self.receipts[scope] = (fingerprint, item_id)
        return item_id

service = WorkItems()
first = service.create("account-a", "approved-action-17", "Review evals")
# Imagine the first response was lost after this commit.
retry = service.create("account-a", "approved-action-17", "Review evals")
second_intent = service.create("account-a", "approved-action-18", "Review evals")
assert first == retry == 1
assert second_intent == 2
assert len(service.items) == 2
try:
    service.create("account-a", "approved-action-17", "Different work")
except ValueError:
    print("conflicting retry rejected")
else:
    raise AssertionError("changed arguments must not reuse a receipt")
print("items:", len(service.items), "retry result:", retry)

Expected output:

Code / output
conflicting retry rejected
items: 2 retry result: 1

What changes in production

The list and dictionary are not durable or safe under concurrent calls. Use a unique database key scoped to the authenticated caller and operation, and atomically commit the effect and receipt when both live in the same database. Concurrent requests must either observe the same completed receipt or receive a documented in-progress outcome.

If the effect happens at a payment provider or another service, a transaction in your own database cannot make both systems atomic. Forward a supported idempotency key to that service and reconcile uncertain outcomes. Do not mark an operation successful merely because you queued it.

Set a documented retention window. Expiring receipts too early can turn a late retry into a new action. Backoff and retry limits reduce load; they do not replace deduplication. Recheck current authorization before exposing stored results. Avoid putting secrets or personal data in keys.

What to test next

Run two concurrent calls with the same key; retry after a process restart; reuse a key with different arguments; retry after the retention window; and simulate the external provider succeeding while your local response handling fails. Measure duplicate effects, not just HTTP success rates.

Source checked 2026-10-07: AWS Builders' Library: Making retries safe with idempotent APIs. The stable caller identifier and atomic effect/receipt guidance comes from this source. The Python scenario is our own AI-assisted teaching example, executed locally; it is not an AWS implementation or a distributed exactly-once guarantee.

0 score

Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.

Verify your sign-in to contribute an answer

Reading is free. Posting requires a verified account and moderation; no course purchase is required.