27 November 2025 / Applied AI / 8 chapters

Following one task across queues and tools

From Self-hosting an AI agent: architecture, observability and recovery

I'd have every component emit structured events with the same identifiers: task_id, run_id, component, event type, timestamp, deployment version and the relevant tool or dependency name. Add tool_call_id, action_id and provider request IDs at those boundaries. Keep user content and credentials out of the ordinary operational fields.

Distributed traces can connect request acceptance, queue wait, worker execution, memory retrieval, model calls and tool operations. OpenTelemetry gives you common conventions for traces, metrics and logs without tying you to one storage backend. Carry the trace context through queued messages and resumed tasks, otherwise the trace ends at dispatch and the slow part looks unrelated.

I wouldn't rely on traces alone, though. Sampling might drop the exact task you're looking at, and a task sitting in the queue doesn't produce new spans. Durable task events give you the full sequence, and metrics show whether a fault is hitting more than one task.

For metrics, I'd pick ones that show blocked work:

  • task counts by current status and task type
  • age of the oldest ready task
  • queue-to-claim delay and active lease count
  • expired leases and repeated attempts
  • dependency errors by stable reason code
  • memory retrieval latency, empty results and timeouts
  • credential expiry windows and authentication failures
  • model and tool call latency, rate limits and uncertain outcomes

Set alerts against service expectations the operators have agreed to support. Queue depth alone can mislead, since a burst of quick jobs might be harmless. Oldest task age, stalled status transitions and a growing pile of unresolved actions usually tell you more.

The operator view should let you piece a task together without trawling several systems by hand: current status, last durable event, dependency state at that time, any active or expired lease, pending approval and evidence for external actions. Link out to restricted detail instead of copying full prompts, tool payloads or memory content into a dashboard lots of people can see.

And test your redaction. Provider errors and tool exceptions often include headers, file paths or bits of the request that the success path never returns. Push representative failure objects through the logging pipeline and check that secrets and private fields are gone before anything reaches log storage.

All articles