22 September 2026 / Agent operations / 8 chapters

Give the operator a procedure they can test

From Recovering an automated workflow after a partial failure

Put the recovery procedure beside the case history or link it directly from the exception queue. It should identify the workflow version it covers and the person responsible for maintaining it. Operators need a current procedure when the alert arrives.

Start with the case lookup and the read-only checks. Then describe the conditions that permit a retry or require escalation. Include the destination account check, the evidence needed to resolve an uncertain write and the result expected after recovery. Use actual commands or screen labels in the internal version, tested against the environment in which people will use them.

AWS's runbook guidance recommends documenting required tools and permissions, error handling and ownership, then asking another team member to validate the procedure. That is a useful test here. Choose someone who has the appropriate role but did not write the recovery code, and watch where they need an explanation that is missing from the document.

Keep the incident discussion separate from the approved procedure. The discussion can contain abandoned attempts and temporary exceptions. The procedure needs the reviewed steps that are safe for the stated conditions.

Give support staff a concise status they can use with the requester. State what is confirmed, what remains unresolved and who has the next action. Avoid presenting an estimated completion time as a promise when recovery depends on a provider investigation. The technical history can retain the detail without forcing every status update to reproduce it.

After a real recovery, check whether the procedure answered the operator's questions. Record any missing step while it is still clear, review the amendment and test the revised instructions before making them the normal path.

All articles