Thrive AI Editorial
Thrive AI Editorial · AI-assisted worked example, not a member conversation.
2026-10-07
Keep the intent stable across transport retries
A timeout tells you that you did not receive a response. It does not prove that the action failed. Retrying the model prompt is also not a reliable way to recover the original tool execution: it may produce a different call or identifier.
Create an operation identifier in your application when the user approves an action. Persist that identifier with the pending work. Use it for every transport retry of that action. A new intentional action gets a new identifier, even if its arguments are identical.
approved intent -> stable operation key -> tool service
-> commit result + key atomically
lost response -> same operation key -> return stored result
Runnable boundary simulation
Python 3, standard library only. This deliberately single-process example demonstrates the contract, not a production database implementation.
import json
class WorkItems:
def __init__(self):
self.items = []
self.receipts = {}
def create(self, caller, key, title):
scope = (caller, key)
fingerprint = json.dumps({"title": title}, sort_keys=True)
if scope in self.receipts:
previous, item_id = self.receipts[scope]
if previous != fingerprint:
raise ValueError("same key, different intent")
return item_id
item_id = len(self.items) + 1
self.items.append({"id": item_id, "title": title})
self.receipts[scope] = (fingerprint, item_id)
return item_id
service = WorkItems()
first = service.create("account-a", "approved-action-17", "Review evals")
# Imagine the first response was lost after this commit.
retry = service.create("account-a", "approved-action-17", "Review evals")
second_intent = service.create("account-a", "approved-action-18", "Review evals")
assert first == retry == 1
assert second_intent == 2
assert len(service.items) == 2
try:
service.create("account-a", "approved-action-17", "Different work")
except ValueError:
print("conflicting retry rejected")
else:
raise AssertionError("changed arguments must not reuse a receipt")
print("items:", len(service.items), "retry result:", retry)
Expected output:
conflicting retry rejected
items: 2 retry result: 1
What changes in production
The list and dictionary are not durable or safe under concurrent calls. Use a unique database key scoped to the authenticated caller and operation, and atomically commit the effect and receipt when both live in the same database. Concurrent requests must either observe the same completed receipt or receive a documented in-progress outcome.
If the effect happens at a payment provider or another service, a transaction in your own database cannot make both systems atomic. Forward a supported idempotency key to that service and reconcile uncertain outcomes. Do not mark an operation successful merely because you queued it.
Set a documented retention window. Expiring receipts too early can turn a late retry into a new action. Backoff and retry limits reduce load; they do not replace deduplication. Recheck current authorization before exposing stored results. Avoid putting secrets or personal data in keys.
What to test next
Run two concurrent calls with the same key; retry after a process restart; reuse a key with different arguments; retry after the retention window; and simulate the external provider succeeding while your local response handling fails. Measure duplicate effects, not just HTTP success rates.
Source checked 2026-10-07: AWS Builders' Library: Making retries safe with idempotent APIs. The stable caller identifier and atomic effect/receipt guidance comes from this source. The Python scenario is our own AI-assisted teaching example, executed locally; it is not an AWS implementation or a distributed exactly-once guarantee.
Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.