22 September 2026 / Agent operations / 8 chapters

Decide how a case picks up again

From Recovering an automated workflow after a partial failure

A case should resume from the last confirmed step whose prerequisites still hold. In the example workflow, a failed notification means you go back to handling the notification. You shouldn't recreate the approved business record just because both actions happened in the same run.

Only let one recovery attempt claim a case at a time. Use an atomic state transition or whatever concurrency control suits your store, and record who or what holds the claim and how it gets released if a worker dies. If claims expire, think about an old worker coming back late. A lease expiring doesn't stop that worker from sending an external request, so destination idempotency or a supported fencing mechanism still has a job to do.

Sort out what kind of error you've got before deciding to retry. A temporary service failure might be worth another go. A rejected field needs corrected input, and a revoked permission needs whoever owns that access. Sending the same invalid request again just makes the queue busier without changing whatever made it fail.

AWS's retry-with-backoff guidance explains using increasing delays for transient failures and the extra load frequent retries cause. Set a bounded retry policy that takes the provider's instructions into account, and check whether the SDK already retries so your application and worker aren't quietly multiplying each other's attempts.

Incoming events can repeat as well. Stripe documents duplicate webhook deliveries, delivery without guaranteed ordering, and manual resends that don't cancel automatic retries. So your recovery path has to cope with an old event turning up while someone's in the middle of fixing the case. Record received events durably and track the work they trigger separately, otherwise a stored receipt can cause unfinished processing to get skipped forever.

When a case hits its attempt limit, send it somewhere visible for human review, with the last error, the actions already confirmed and the next check that's allowed. I wouldn't leave it sitting there labelled "processing" after every worker has stopped.

All articles