Article chapter 04 of 08
Build health checks around task paths
From Self-hosting an AI agent: architecture, observability and recovery
Use separate startup, liveness and readiness checks. Startup tells the supervisor that a component has loaded configuration and completed required initialisation. Liveness answers whether restarting that component may help. Readiness answers whether it should receive new work now. Combining all three in one deep endpoint can cause a brief database fault to restart every healthy process at once.
A worker may be live but unready because it cannot claim a task, reach a required model adapter or load its tool registry. The entry service may remain ready to show existing task state while refusing new tasks that depend on an unavailable queue. Readiness should follow actual routing decisions rather than report every dependency as equally fatal.
Dependency checks need to exercise the operation the agent uses. Opening a TCP connection to the memory database does not prove that the application can authenticate, read the current schema or retrieve a permitted record. A queue check should confirm publishing and consumption through a dedicated probe path. Credential checks should inspect expiry and, where safe, perform a low-cost authorised read. Secret values never belong in the health response.
Keep frequent probes bounded. A health endpoint that runs a model generation, broad search and several external API calls can create load and provider cost during the incident it is meant to diagnose. Use cheap local checks for routing, then run synthetic tasks on a slower schedule.
A useful synthetic task follows a production-shaped path with controlled data. It creates a task, dispatches it, retrieves a known non-sensitive memory, calls a stub or harmless read-only tool and verifies the recorded result. Run separate probes for scheduled work and approval resumption if those paths use different machinery. Tag synthetic records so retention and alerts can handle them without hiding their failures.
Expose dependency state with its check time and reason. A memory: degraded result should also show the last successful retrieval and current timeout. memory: false does not support a decision. Show stale checks as unknown. A probe process that stopped reporting should never leave yesterday's green result on the screen.