Article chapter 02 of 08
Map the components and their state
From Self-hosting an AI agent: architecture, observability and recovery
A small installation can run on one host without collapsing everything into one application process. Separate components by responsibility, durable state and restart behaviour, even if containers or system services place them on the same machine.
The entry service authenticates requests, validates their basic shape and creates a durable task record. A scheduler finds eligible work. Workers claim a task, assemble context, call a model and invoke permitted tools. A tool gateway validates arguments, applies policy and holds access to external credentials. Memory retrieval may involve a relational store, object storage and a search or vector index. An event store records what happened. The operator interface reads task and dependency state without becoming the source of truth for either.
The exact products matter less than explicit ownership. Record which component owns each transition and which store is authoritative. A queue can say that a message was delivered while the task database still says queued. A search index can contain a memory whose source record was deleted. An interface cache can show running after a worker lease expired. Recovery depends on knowing which record wins and how derived state is rebuilt.
Keep model-provider and tool traffic behind narrow adapters. The worker should receive a typed result such as rate_limited, credential_expired, policy_denied or outcome_uncertain. Raw exceptions encourage accidental retries and make provider changes leak through the task state model. Adapters also create one place to apply timeouts, redact logs and attach correlation identifiers.
Document the network boundaries as deployed. A container name in a Compose file does not prove that the service is unreachable from outside the host. Check published ports, firewall rules, reverse proxy routes and administration interfaces. Databases, queues, tracing collectors and model gateways usually need private connectivity only. If remote administration is required, put it behind authenticated access rather than exposing each component separately.
Configuration should identify its version. Store deployment configuration in a controlled repository, keep secrets in a secret store or protected runtime environment, and record the application and migration versions with each release. A restored database paired with unrecorded configuration changes is a common way to turn recovery into guesswork.