27 November 2025 / Applied AI / 8 chapters

Bound load, retries and degraded operation

From Self-hosting an AI agent: architecture, observability and recovery

Self-hosted capacity is finite, even when model inference happens elsewhere. A flood of accepted tasks can exhaust database connections, fill a disk with events or hold every worker open on a slow dependency. Admission and scheduling limits should preserve enough capacity for status checks, cancellation and recovery work.

Set concurrency by workload class. Long document processing, interactive requests and scheduled maintenance do not need to share one undifferentiated pool. Reserve small amounts of capacity where an interactive or recovery task must remain responsive. Place hard limits on task runtime, tool calls, retrieved content, generated output and queued payload size.

Timeouts need to fit inside a task budget. If a worker allows four sequential dependencies to each consume the whole task timeout, cancellation and lease expiry become unpredictable. Pass a remaining deadline through adapters and stop starting work that cannot reasonably finish inside it.

Retry only faults that another attempt may change. Connection refusal before a request was sent may be retryable with bounded backoff. Invalid arguments, policy denial and expired approval need a different transition. A tool timeout after sending a write creates an uncertain outcome and should move to reconciliation.

Circuit breakers can stop repeated calls to a failing dependency while allowing occasional probes for recovery. During that period, the scheduler can hold tasks whose contract requires the dependency. Other task types may continue through an explicitly defined reduced path. The result should state what was unavailable, and the task record should keep the degraded-mode decision.

Watch local resources alongside application metrics: disk capacity and inode use, memory pressure, CPU saturation, database connections, queue storage and backup age. Set retention and compaction rules before logs, traces, model artefacts and task events compete with the database for the last free space. Keep the durable audit material required for recovery, and expire diagnostic copies according to their own policy.

All articles