22 September 2026 / Agent operations / 8 chapters

A procedure the operator can actually test

From Recovering an automated workflow after a partial failure

Put the recovery procedure next to the case history, or link to it straight from the exception queue. It should say which workflow version it covers and who's responsible for keeping it up to date, because when the alert comes in, the operator needs the current version.

Start with the case lookup and the read-only checks, then say when a retry is allowed and when it needs escalating. Include the destination account check, the evidence needed to settle an uncertain write and what the result should look like afterwards. The internal version should use the actual commands or screen labels, tested in the environment people will be using.

AWS's runbook guidance recommends documenting the tools and permissions needed, error handling and ownership, and then getting another team member to validate the procedure. That's a good test here. Pick someone with the right role who didn't write the recovery code, and watch where they need an explanation the document doesn't give them.

Keep the incident discussion (abandoned attempts, temporary exceptions) separate from the approved procedure, which should only have reviewed steps that are safe under the stated conditions.

Support staff need a short status they can pass on to the requester: what's confirmed, what's still open and who's doing the next thing. Don't present an estimated completion time as a promise when recovery depends on a provider investigating.

After a real recovery, check whether the procedure answered the operator's questions. Write down any missing step while it's still fresh, get the change reviewed and test the revised instructions before they become the normal path.

All articles