19 December 2024 / Applied AI / 8 chapters

Work out detection and recovery before you automate anything

From Production readiness for applied AI

Once a workflow is in production you need to be able to tell accepted work, rejected work, incomplete work and work whose final state is unknown apart. If you can't, support staff end up piecing events together from chat text and timestamps while the customer waits.

Give every run a stable identifier and every external action its own operation identifier. Then store the request identity, workflow version, source references, model and prompt version, proposed output, validation result, approval event, tool request and confirmed response. Sensitive input should follow the system's retention rules, but the record still needs enough references to find the authorised evidence when someone's investigating.

Use states that match what the software can actually prove. proposed, approved, submitted, confirmed, failed and cancelled tell you far more than a single done flag. If an external call times out, the operation might need to sit in unknown until someone checks the receiving system. Marking it failed and retrying automatically can duplicate an action that went through fine before the response got lost.

Recovery should be a supported path, either in the application or in the operating procedure. For each state, decide which transitions are allowed and who can trigger them. A reviewer might edit a proposal and resubmit it. An operator might reconcile an unknown operation against the external system. A confirmed action might need a compensating transaction, because you can't just delete it.

I'd test recovery as seriously as the main path. Interrupt the request after the model returns, after approval, in the middle of a tool call, and after the external action succeeds but before the local confirmation is saved. Restart the worker or the application and look at the state it's left in. Either it carries on safely or it leaves a visible item a person can resolve.

A recovery test passes when the operator can answer these from the evidence they've got:

  • What was the workflow trying to do?
  • Which step finished?
  • Did an external system change?
  • What's safe to do now?

Don't bury unresolved work in a general error count. Put unknown and failed operations into a queue showing age, owner, last attempt and the next allowed action. Make sure somebody watches it and has the tools to clear it.

All articles