22 September 2026 / Agent operations / 8 chapters

Recovering an automated workflow after a partial failure

When an automated workflow stops halfway through, some of its changes may already exist in other systems. Before restarting it, establish what completed and which actions can be repeated safely.

A row of industrial control cabinets with the centre cabinet open for inspection.
Generated image: Villar David editorial
Read the Introduction

An automated workflow submits an approved record to another system, waits for a response and times out. The operator sees a failed run. The destination may already contain the record. Starting again could create a duplicate, while leaving it alone could leave genuine work unfinished.

OpenAI's Agents API announcement this month describes agents that can keep working across long sessions, use tools and delegate parts of a task to subagents. Once that work can continue for hours or days, the recovery path becomes part of the design. More steps can finish before the failure reaches an operator.

This guide uses a hypothetical document-processing workflow: receive a document, prepare a proposed record, obtain approval, write to an external system and notify the requester. An AI agent might prepare the proposal or help investigate an exception. The recovery design still needs to account for what each external action actually did.

The aim is to give the next operator enough information to resume a case without guessing. The record fields and checks below are proposed design choices. They need adapting to the destination's API, the organisation's approval rules and the consequences of a duplicate or missing action.

Chapter 1: Record each external action separately

All articles