Article chapter 07 of 08
Restore authoritative state in a known order
From Self-hosting an AI agent: architecture, observability and recovery
Backups are useful after they have been restored under the same constraints as the running system. List each durable store and decide whether it is authoritative, derived or disposable. The task database and event records may be authoritative. A vector index may be rebuildable from approved memory records and source objects. A local model cache can usually be repopulated.
Define recovery objectives for each store based on consequences. Losing recent task events may leave external actions unresolved. Rebuilding a search index may delay memory retrieval without losing its source. Those two stores should not inherit the same backup schedule merely because they share a host.
Coordinate snapshots where records span stores. A database backup that references object versions absent from the object-store snapshot can restore cleanly and still produce broken tasks. Where atomic snapshots are unavailable, use versioned objects, a recorded backup boundary and reconciliation after restore.
A practical restoration order is documented rather than improvised. Restore secrets infrastructure or credential references, then authoritative databases and object data. Apply the application version and schema migrations that match the backup. Start queues and schedulers in a paused state. Rebuild derived indexes, check referential integrity and inspect tasks that were running or unresolved at the backup boundary. Only then admit new work.
Do not restore old secret values merely to make the recovered system match. Rotate exposed or uncertain credentials and update references through the supported path. Test that agents cannot fall back to a credential left in an old environment file, container layer or backup bundle.
Keep deployment definitions, migration scripts, tool schemas, policy versions and recovery procedures outside the failed host. A copy that exists only on the machine being recovered is part of the incident. The recovery operator should be able to identify the expected versions without relying on shell history.
Run restore tests into an isolated environment. Check that a known completed task remains complete, a pending task can be resumed, source references still open, derived memory can be rebuilt and an unresolved action remains blocked until reconciliation. Record restore duration and any manual steps as observations from the drill, then update the procedure.