Article chapter 08 of 08
Practising the failures that leave work half done
From Self-hosting an AI agent: architecture, observability and recovery
For recovery you need to break things on purpose around the points where state gets saved, using controlled tasks and reversible or stubbed external actions so a drill can't touch private or production records.
I'd start with the failures that can leave the interface looking healthy:
- stop the queue consumer while the entry service keeps accepting requests
- make memory retrieval time out while its database port stays open
- expire a tool credential between the agent proposing an action and running it
- stop a worker after it sends an external request but before it saves the response
- fill the tracing or log destination and check that task execution doesn't quietly lose the audit events it needs
- restart the scheduler while tasks are waiting for approval
- restore a backup without the derived indexes
For each drill, look at what the user sees, the durable task state, the alerts, the dependency page and what the operator is meant to do. Nothing should be reported complete without evidence from the destination. An empty memory result should look different from unavailable memory, an expired lease shouldn't trigger an unsafe replay, and held tasks should only resume once their dependency or approval condition is actually met.
Map recovery actions to reason codes. A credential failure might need rotation and a fresh authorisation check. Queue delay might need the consumer fixed and then a controlled release of the backlog. An uncertain external write needs reconciling against the destination. A missing derived index needs a rebuild while the affected task types stay paused or clearly marked degraded.
Before release I'd want to see one host restart and one backup restore done with someone watching, starting from queued, running, approval-waiting and unresolved test tasks. Afterwards, account for every task and every controlled external effect (for the restore, at the recorded backup boundary).
Then take one production-shaped synthetic task and stop its worker right after a tool request goes out. Can the system show the unresolved action, stop a duplicate, reconcile with the destination and resume the task from its recorded state? If it can, it's ready for a more limited first deployment. If any step relies on someone remembering what happened from a terminal window, add the missing event or procedure before you give the agent real authority.
An agent whose interface loads fine while the queue, memory or a tool credential behind it has stopped working is the case this whole article has been about. Self-hosting it well mostly comes down to the system reporting on the path each task actually takes, keeping enough durable state that a restart or restore doesn't lose track of half-done work, and practising the failures that cause those half-broken states. If I were setting one up now, I'd write down the task types and their dependencies first, then build the readiness checks, task events and restore order around them, and run the drills above before trusting the green status on the dashboard.