27 November 2025 / Applied AI / 8 chapters

Limits on load, retries and running degraded

From Self-hosting an AI agent: architecture, observability and recovery

Even with model inference happening elsewhere, your host has limits. A flood of accepted tasks can use up database connections, fill a disk with events or tie up every worker on one slow dependency, so set admission and scheduling limits that leave room for status checks, cancellation and recovery work.

Set concurrency by workload. Long document processing, interactive requests and scheduled maintenance don't need to share one big pool, and you can hold back a little capacity where an interactive or recovery task has to stay responsive. Put hard limits on task runtime, tool calls, retrieved content, generated output and queued payload size.

Timeouts have to fit inside the task's overall budget. If a worker lets four sequential dependencies each use the whole task timeout, cancellation and lease expiry get unpredictable. Pass the remaining deadline down through the adapters and don't start work that can't reasonably finish in the time left.

Only retry faults that another attempt might change. A refused connection before the request went out may be fine to retry with bounded backoff. Invalid arguments, policy denial and an expired approval need a different transition. A tool timeout after a write was sent is an uncertain outcome, so that goes to reconciliation.

Circuit breakers can stop repeated calls to a failing dependency while letting the odd probe through to see if it's back. While one's open, the scheduler can hold tasks that need that dependency, and other task types can keep going through a reduced path you've defined ahead of time. The result should say what wasn't available, and the task record should note that it ran degraded.

Watch local resources alongside application metrics: disk space and inode use, memory pressure, CPU saturation, database connections, queue storage and the age of the last backup. Set retention and compaction rules before logs, traces, model artefacts and task events end up fighting the database for the last bit of free space. Expire diagnostic copies on their own schedule, separately from the audit material you need for recovery.

All articles