22 September 2026 / Agent operations / 8 chapters

Control how a case resumes

From Recovering an automated workflow after a partial failure

Resume from the last confirmed step whose prerequisites still hold. In the example workflow, a failed notification should resume notification handling. It should not recreate the approved business record merely because both actions belong to the same run.

Allow one recovery attempt to claim the case at a time. Use an atomic state transition or another concurrency control suitable for the store. Record who or what holds the claim and how it can be recovered after a worker dies. If claims expire, account for an old worker returning late: a lease expiry on its own does not prevent that worker from submitting an external request. Destination idempotency or a supported fencing mechanism still has a job to do.

Classify errors before deciding to retry. A temporary service failure may justify another attempt. A rejected field needs corrected input. A revoked permission needs the appropriate owner. Repeating the same invalid request makes the queue busier without changing the condition that caused it to fail.

AWS's retry-with-backoff guidance explains the use of increasing delays for transient failures and the extra load frequent retries can cause. Set a bounded policy for the operation, taking account of provider instructions. Check whether the SDK already retries so the application and worker don't unknowingly multiply each other's attempts.

Incoming events can also repeat. Stripe documents duplicate webhook deliveries, delivery without guaranteed ordering, and manual resends that do not cancel automatic retries. Your recovery path therefore needs to tolerate an old event arriving while someone is fixing the case. Record received events durably, then track the work they trigger separately. A stored receipt must not cause unfinished processing to be skipped forever.

When the attempt limit is reached, give the case a visible destination for human review. Include the last error, the actions already confirmed and the next permitted check. Avoid leaving it indefinitely labelled "processing" after every worker has stopped.

All articles