27 August 2026 / Applied AI / 8 chapters

Price failed runs and correction

From Budgeting and governing an AI workflow after the prototype

A failed run may trigger a repeated model call, reconciliation of an external action, partial-state recovery, user notification and support work. Price the recovery path as part of the run.

List failure modes by stage and decide the recovery path for each. A provider timeout before any action may be safe to retry. A timeout after sending an instruction to another system may leave an uncertain result that needs reconciliation. A malformed structured response can be retried within a limit, while a policy conflict may need a person.

For each path, record:

  • detection mechanism;
  • retry rule and maximum attempts;
  • idempotency or duplicate protection;
  • state preserved for recovery;
  • person or service that owns the exception;
  • user communication;
  • correction or reversal process;
  • data retained as evidence;
  • expected handling cost.

Put limits on retries. An unbounded loop can turn a provider problem into a large bill and a backlog. Use a circuit breaker or operational pause when repeated failures cross an agreed threshold. The threshold should be based on workflow consequence and normal variation, then tested in a controlled environment.

Budget for wrong accepted results as well as technical failures. Correction may require finding affected records, restoring prior values, contacting users and explaining the decision. Some actions cannot be fully reversed. Those workflows need stronger approval controls and a cost assumption for remediation before launch.

Maintain a small reserve for investigation and correction rather than pretending every cost is predictable. Set it from the risk review and revise it as evidence accumulates. Keep it visible; a generic contingency line gives little help when deciding which failure mode needs prevention.

All articles