Article chapter 01 of 08
When the agent looks fine and isn't
From Self-hosting an AI agent: architecture, observability and recovery
You open the agent's interface, it loads, takes your request and shows a working indicator. Meanwhile the queue behind it stopped dispatching jobs an hour ago. On another install, model calls still work but memory retrieval keeps timing out. On a third, the agent prepares an action and then finds its tool credential expired overnight. The application process is alive in all three, and the task is either stuck or running without the context it needs.
A process check only tells you one process answered. If you're self-hosting an agent, you want the system to report on the path a task takes through storage, scheduling, model calls and tools, including the in-between states where some work is still safe to run.
Before drawing any deployment diagram, I'd write down a handful of task types. A read-only research task, a scheduled task that uses durable memory and a task that waits for approval before changing an external record all take different paths. For each one, write down:
- where the request first gets saved somewhere durable
- which stores and services it has to reach
- which credentials it uses
- where it can sit and wait
- which effects can leave the host
- what counts as evidence that it's done
You'll find the weak dependencies pretty quickly doing this. If memory is optional for a task type, the agent can carry on and say memory wasn't available. If the task needs current memory, it should stop with a dependency error. What you don't want is a failed retrieval quietly treated as an empty result, because then an outage looks like "there's nothing relevant in memory".
Give every dependency an owner and a failure policy: does new work get rejected, queued, run in a reduced mode or sent to an operator? I'd avoid a single global healthy flag. A status page saying scheduled work is delayed is far more useful than a green badge next to an interface that can't finish its jobs.