Article chapter 03 of 08
Keeping task state through a worker restart
From Self-hosting an AI agent: architecture, observability and recovery
A worker can stop after nearly any step: after the model responds, while a tool's running, after an external request has gone out, or right before it saves the result. Whatever's stored durably has to tell the next worker what it can safely do.
Create a task_id when the request is accepted and a run_id for each attempt. Store the objective, authenticated actor, scope, creation time, current status and the class of dependency it needs. A status set might be queued, running, waiting_for_approval, waiting_for_dependency, completed, failed, cancelled and unresolved. Keep the whole transition history as well as the latest label.
Workers should claim work with a lease that has an owner, a token and an expiry, and only the current lease holder can mark the task complete. If a worker disappears, the scheduler waits for the lease to expire and then looks at the task's last durable step. An expired lease doesn't automatically mean it's safe to run again.
External actions need a stable identity. Create the action record before sending the request, and use its ID as an idempotency key if the destination supports one. After a timeout, reconcile using that key or the destination's object reference. Sending a fresh request can give you a duplicate email, job or record when the first one actually went through.
Waiting tasks shouldn't eat worker capacity. A task waiting on approval or a broken dependency should give up its lease and keep a durable wake-up condition, so the scheduler can resume it when an approval event arrives, a dependency changes state or a review time comes around. Polling loops inside workers waste capacity and vanish on restart.
Cancellation needs state too. Stop new model and tool calls, mark requests that might still be in flight and reconcile them. That way cancellation_requested looks different from a task where every external effect is accounted for.