27 November 2025 / Applied AI / 8 chapters

Mapping the pieces and who owns what

From Self-hosting an AI agent: architecture, observability and recovery

A small install can live on one host and still keep components split by responsibility, durable state and restart behaviour. Here's roughly how I'd lay it out. The entry service authenticates requests, checks their basic shape and creates a durable task record. A scheduler finds work that's ready. Workers claim a task, pull together context, call a model and use whatever tools they're allowed. A tool gateway checks arguments, applies policy and holds the external credentials. Memory retrieval might involve a relational store, object storage and a search or vector index. An event store records what happened. The operator interface reads task and dependency state, but it isn't the source of truth for either.

Which products you pick matters much less than being clear about ownership. Write down which component owns each transition and which store wins when they disagree. A queue can say a message was delivered while the task database still says queued. A search index can hold a memory whose source record was deleted. An interface cache can keep showing running after a worker's lease expired. In a recovery you need to know that, and how to rebuild the derived stuff.

I'd keep model-provider and tool traffic behind narrow adapters, so the worker gets back a typed result like rate_limited, credential_expired, policy_denied or outcome_uncertain. Raw exceptions invite accidental retries and let provider changes leak into your task states. Adapters also give you one place for timeouts, log redaction and correlation IDs.

Check the network boundaries as actually deployed. A container name in a Compose file doesn't mean the service is unreachable from outside the host, so look at published ports, firewall rules, reverse proxy routes and admin interfaces. Databases, queues, tracing collectors and model gateways usually only need private connectivity, and remote admin belongs behind authenticated access.

Version your configuration too. Keep deployment config in a controlled repo, secrets in a secret store or protected runtime environment, and record the application and migration versions with each release. Restoring a database next to config changes nobody wrote down turns recovery into guesswork.

All articles