Article chapter 03 of 08
Make a retry count as the same operation
From Recovering an automated workflow after a partial failure
If the API supports idempotency, use its documented mechanism for any operation you might need to retry. The idea is that sending the same request again has no extra effect. How well that protects you depends on the provider's actual contract, including how the identifier is scoped and how long it lasts.
Create the operation identifier before the first attempt and hang on to it. A retry of the same operation has to reuse that identifier. If you generate a new one after every timeout, each attempt can look like a separate request to the destination.
It also needs to be specific enough to tell legitimate actions apart. Creating an external record and later updating it are two different operations, while two workers trying the same approved creation should land on the same operation even though their execution identifiers differ. Put the tenant or account into your local uniqueness rules so unrelated customers can't collide.
Featonby's article also covers what happens when a request reuses an identifier but the parameters have changed. Keep the approved payload, or a protected reference that can reproduce it, so a retry doesn't quietly send different content under the old identifier. If the requester corrects the document, decide on purpose whether that's a new operation and whether the old one has to be reconciled first.
If you control the receiving service, do the deduplication atomically with the local change it's protecting. A separate "have I seen this?" query followed by an unprotected insert lets two concurrent workers pass the check at the same time. A unique constraint and the right transaction can enforce the rule inside that database, though that guarantee doesn't automatically carry over to another service you call afterwards.
And if the destination doesn't have usable idempotency support, you still need the uncertain-result path. A local lock can cut down on concurrent attempts, but it can't prove a remote request failed after the lock holder lost contact. Some cases will need someone to reconcile them by hand before another write is allowed.