Blog
Operating local AI infrastructure after the migration
A successful migration is followed by less glamorous work: services starting in the right order, data paths staying stable, backups being restorable and failures being visible. I am treating those operating details as part of the AI product because continuity disappears when the infrastructure is opaque.

A migrated system can look healthy while the person who performed the move still has shells open, environment variables loaded and services started by hand. The next reboot removes that temporary scaffolding. It is the quickest way to discover whether the machine can return to service without remembering a private sequence of commands. I want each component managed by the operating system or the container runtime, with a declared restart policy and a known service user. The agent runner should not begin accepting work until its database, memory store and required local APIs are ready. Use a readiness check that performs a small real operation before the runner accepts work. Startup dependencies also need time limits. If the memory service does not become ready, the agent should stay unavailable or enter a clearly reduced mode. Silently starting without memory would produce work under different conditions while the dashboard still showed a green process. After a migration, paths that once happened to exist become dependencies. Prompts refer to supporting files. Tools expect working directories. Backup scripts target database volumes. A change to a mount point can break all of them without changing application code. I am keeping authoritative data, configuration, generated indexes, caches and logs in separate locations. The separation makes the backup boundary explicit and prevents a cleanup of rebuildable files from deleting source records or active configuration.
Service configuration should use the same documented paths as recovery instructions. If a container maps a host directory, the mapping belongs in versioned configuration rather than an ad hoc launch command. Permissions need to be tested under the service account, not under my interactive user. A useful path check creates and reads a disposable record through the running service, confirms where it landed, then removes it through the same interface. This catches cases where a service is writing to an unexpected empty volume while the real data sits untouched elsewhere. A backup job that exits successfully may still produce an incomplete or unusable artefact. Database files copied while a service is writing can be inconsistent. An encrypted archive can become inaccessible when its key or restore command is missing. Retention can quietly keep many copies of the same corruption. I want backup jobs to record the source, start and finish times, artefact size, checksum, encryption method and retention action. Use those details to detect obvious changes, then restore the artefact into an isolated location. For the local AI stack, the restore test should cover the records that cannot be recreated: durable memory, source links, agent configuration and activity history. Derived indexes can be rebuilt if their inputs and build settings are retained. The rebuild itself needs testing because a dependency or model version may no longer match the saved configuration.
Run a monthly isolated restore rather than relying on checks that archive files exist. The restored services should answer a few fixed queries, including one that follows a memory record back to its source. Once the test passes, the temporary environment can be removed and the result kept with the backup log. A single endpoint that returns 200 OK often reports that the web process can answer HTTP. It says nothing about the job queue, database, disk space or model endpoint. I prefer small checks at each layer, then one synthetic task that crosses the main path. The component checks can cover database connectivity, memory retrieval, queue access, available disk space and age of the most recent backup. The synthetic task should be bounded and harmless. It might submit a read-only job, retrieve a known test record and write an activity event to a test namespace. Each check needs a timeout and an owner. Without a timeout, a stuck dependency can consume every worker. Without an owner, an alert becomes background noise. I also want the failure response written down: restart a service, pause new jobs, switch to read-only mode or escalate for inspection.
Liveness shows that a process is alive. Readiness shows whether it can do useful work, which is the signal schedulers and reverse proxies should use before sending jobs. Local infrastructure spreads one job across several processes. The scheduler accepts it, a worker invokes tools, the memory layer performs retrieval and a database records results. Timestamp searches are painful when clocks or retries overlap. I am carrying a task identifier and run identifier through those components. The task identifies the requested piece of work. A retry creates another run under the same task. Log entries can then show which attempt called which tool, what state transition occurred and where an error started. Ordinary logs should avoid prompt contents, credentials and private source text. Structured fields can carry identifiers, component names, durations, status and error classes without copying sensitive payloads. When deeper inspection is needed, the task record can point to access-controlled detail. Log rotation is part of the design. An always-running local service can fill a disk with debug output and take down the database it was meant to support. I want size or time limits, retention and a disk-space alert tested before verbose logging is enabled for an investigation.
The migration is complete, but the new host will keep changing through package updates, model changes, credential rotation and growing data. I need a small operating routine that notices drift without turning maintenance into a full-time project. My weekly check is short: review failed and retried jobs, confirm backup recency, inspect disk growth, check expiring credentials and look for services that restarted unexpectedly. After configuration changes, I rerun the synthetic task. After updates that affect storage, I run a backup and restore check before removing the previous version. I am also keeping a rebuild note for the host itself. It should list the base system, service definitions, required packages, data restore order and acceptance tests. A second operator should be able to use the note without relying on facts held only in my head. The next concrete check is a cold reboot followed by the synthetic read-only job. I want the run record to show every required service became ready in order, the known memory record was retrieved and the result was written without an interactive fix.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.