Article chapter 08 of 08
Practise the failures that leave ambiguous work
From Self-hosting an AI agent: architecture, observability and recovery
A normal end-to-end test proves the clean path. Recovery needs deliberate interruption around durable boundaries. Use controlled tasks and reversible or stubbed external actions so the drill cannot affect private or production records.
Start with failures that can leave the interface looking healthy:
- stop the queue consumer while the entry service continues accepting requests;
- make memory retrieval time out while its database port remains open;
- expire a tool credential between proposal and execution;
- stop a worker after it sends an external request but before it stores the response;
- fill the tracing or log destination and confirm task execution does not silently lose required audit events;
- restart the scheduler while tasks are waiting for approval;
- restore a backup with derived indexes absent.
For each drill, inspect what the user sees, the durable task state, alerts, dependency page and operator action. Confirm that the system does not report completion without destination evidence. Check that an empty memory result is distinct from unavailable memory, an expired lease does not cause an unsafe replay, and held tasks resume only after their dependency or approval condition is satisfied.
Assign recovery actions by reason code. A credential failure may require rotation and a new authorisation check. Queue delay may require consumer repair followed by controlled release. An uncertain external write needs destination reconciliation. A missing derived index calls for rebuild while affected task types remain paused or clearly degraded.
The release check should include one witnessed host restart and one restored backup. Begin with queued, running, approval-waiting and unresolved test tasks. After restart, account for every task and every controlled external effect. After restore, repeat the check at the recorded backup boundary.
Take one production-shaped synthetic task and stop its worker immediately after a tool request leaves. If the system can show the unresolved action, prevent a duplicate, reconcile the destination and resume the task through its recorded state, the recovery path is ready for a more limited first deployment. If any step depends on remembering what happened from a terminal window, add the missing event or procedure before giving the agent real authority.