Article chapter 07 of 08
Restoring state in the right order
From Self-hosting an AI agent: architecture, observability and recovery
A backup is only useful once you've restored it under the same constraints as the running system. List every durable store and decide whether it's authoritative, derived or disposable. The task database and event records are probably authoritative. A vector index can usually be rebuilt from approved memory records and source objects. A local model cache can normally just be repopulated.
Set recovery objectives per store based on what you'd lose. Losing recent task events might leave external actions unresolved, while rebuilding a search index only slows memory retrieval for a while because the source is still there. Those two stores shouldn't get the same backup schedule just because they share a host.
Coordinate snapshots where records span stores. A database backup that points at object versions missing from the object-store snapshot can restore without a single error and still give you broken tasks. If you can't get atomic snapshots, use versioned objects, record the backup boundary and reconcile after the restore.
Write the restore order down so nobody's improvising it on the night:
- Restore the secrets infrastructure or credential references.
- Restore the authoritative databases and object data.
- Apply the application version and schema migrations that match the backup.
- Start queues and schedulers paused.
- Rebuild derived indexes, check referential integrity and look at any task that was running or unresolved at the backup boundary.
- Then let new work in.
Don't restore old secret values just so the recovered system matches what it was. Rotate any credential that was exposed or that you're unsure about, update references through the normal path, and test that agents can't fall back to a credential left in an old environment file, container layer or backup bundle.
Keep deployment definitions, migration scripts, tool schemas, policy versions and recovery procedures off the host that failed. If the only copy is on the machine you're recovering, it's part of the incident.
Run restore tests into an isolated environment. Check that a known completed task is still complete, a pending task can be resumed, source references still open, derived memory can be rebuilt and an unresolved action stays blocked until it's reconciled. Write down how long the restore took and any manual steps, then update the procedure.