Article chapter 02 of 08
Design detection and recovery before automation
A production workflow needs a way to distinguish accepted work, rejected work, incomplete work and work whose final state is unknown. Without those states, support staff must reconstruct events from chat text and timestamps while the customer waits.
Give every run a stable identifier and every external action its own operation identifier. Store the request identity, workflow version, source references, model and prompt version, proposed output, validation result, approval event, tool request and confirmed response. Sensitive input should follow the system's retention rules. The record still needs enough references to locate authorised evidence during an investigation.
Define states that match what the software can prove. proposed, approved, submitted, confirmed, failed and cancelled are clearer than a single done flag. An external timeout may leave an operation in unknown until the receiving system is checked. Marking it failed and automatically trying again can duplicate an action that succeeded before the response was lost.
Recovery should be a supported path in the application or operating procedure. For each state, decide which transitions are permitted and who can initiate them. A reviewer might edit and resubmit a proposal. An operator might reconcile an unknown operation against the external system. A confirmed action may require a compensating transaction rather than deletion.
Test recovery with the same seriousness as the main path. Interrupt the request after the model returns, after approval, during a tool call and after the external action succeeds but before local confirmation is stored. Restart the worker or application and inspect the resulting state. The system should either continue safely or leave a visible item that a person can resolve.
A recovery test passes when the operator can answer four questions from available evidence:
- What was the workflow trying to do?
- Which step completed?
- Did an external system change?
- What action is safe now?
Do not hide unresolved work inside a general error count. Put unknown and failed operations into a queue with age, owner, last attempt and next permitted action. Assign somebody to watch that queue and give them access to the tools needed to clear it.