Article chapter 04 of 08
Health checks that follow how tasks actually run
From Self-hosting an AI agent: architecture, observability and recovery
I'd use separate startup, liveness and readiness checks. Startup tells the supervisor a component has loaded its config and finished initialising. Liveness answers "would restarting this help?" Readiness answers "should this get new work right now?" Roll all three into one deep endpoint and a brief database blip can restart every healthy process at once.
A worker can be live but not ready because it can't claim a task, reach a required model adapter or load its tool registry. The entry service might stay ready to show existing task state while refusing new tasks that need a queue that's down. Not every dependency is equally fatal, and readiness should reflect that.
Dependency checks should exercise the operation the agent actually uses. A TCP connection to the memory database doesn't tell you the app can authenticate, read the current schema or pull back a record it's allowed to see. A queue check should publish and consume through a dedicated probe path. Credential checks should look at expiry and, where it's safe, do a cheap authorised read. Secret values never go in the health response.
Keep the frequent probes cheap. A health endpoint that runs a model generation, a broad search and a few external API calls adds load and provider cost in the middle of the incident you're diagnosing. Use cheap local checks for routing and run synthetic tasks on a slower schedule.
A good synthetic task follows a production-shaped path with controlled data: create a task, dispatch it, retrieve a known non-sensitive memory, call a stub or harmless read-only tool and check the recorded result. If scheduled work and approval resumption use different machinery, give them their own probes. Tag synthetic records so retention and alerting can tell them apart, without hiding their failures.
Show dependency state with when it was checked and why. memory: degraded should come with the last successful retrieval and the current timeout; memory: false helps nobody. Show stale checks as unknown, so a probe that stopped reporting doesn't leave yesterday's green result on screen.