22 September 2026 / Agent operations / 8 chapters

Testing the gaps between steps that worked

From Recovering an automated workflow after a partial failure

Use a test environment with controlled inputs and destinations, and interrupt the workflow at the points where something actually happens. Then judge the recovery by looking at stored state and destination records, since a worker returning successfully doesn't tell you what's in the destination.

For the example workflow, I'd want to try:

  • The destination accepts a write, but the response gets lost.
  • The response arrives, but the worker stops before saving completion locally.
  • A duplicate event arrives during recovery.
  • Two workers try to claim the same case.
  • The approved input changes while a case is waiting.
  • A notification fails after the external record has been confirmed.
  • The destination lookup is down or temporarily incomplete.
  • A compensating action fails after it's been authorised.

For each one, write down the destination records you expect and the work that should be left before you run the test. Look for duplicates as well as missing changes. Check that an uncertain result stays visible, that an unapproved change can't go ahead, and that another operator can recover the case from the saved evidence.

Try the slow path too. An immediate retry might fall inside a provider's deduplication window when a manual recovery next week won't. Use the provider's documented limits to decide what your own records need to keep and when you'll need a fresh lookup or a fresh approval.

Then compare the recovered result with the source and the approved proposal. The note on checking an integration after the job says it finished goes through that read-only comparison. Any differences you can't resolve should stay assigned to someone and visible, even after the execution itself has finished.

Before you turn recovery on for live work, get another authorised operator to run the lost-response test from the procedure. Ask them to find the existing destination record, complete only the outstanding steps and show you the final comparison. Keep what they found with the workflow release record, including anything that still needed manual investigation.

Which brings me back to the run that timed out after sending the approved record. Whether it's safe to restart comes down to what the destination actually did, and you can only answer that if each external action was recorded with its own identifier before it was sent and the uncertain result was left open until someone checked. I'd start with the read-only lookup in the right destination account, confirm which steps really finished, and only then resume from the last confirmed step using the same operation identifier. If the lost-response test can't be recovered that way from the saved evidence, I'd fix the records and the procedure before letting recovery anywhere near live work.

All articles