Article chapter 02 of 09
Give every task an ID that survives retries
Every event has to belong to a stable task. A browser session, a model response ID or a background process ID won't do, because one task can easily span all three.
Create a task_id when the work is accepted. Store who asked for it, the authenticated tenant or workspace, the task type, a bounded objective, when it was created, its current status and the parent task if another workflow handed it off. If the agent kicks off a subtask, give that its own ID and link it to the parent so the two sequences don't get mixed together.
Each execution attempt gets a run_id as well. The task_id stays the same when you retry under a different model, resume after an approval or let another worker pick it up, and each of those technical attempts gets its own run_id. I'd only keep an attempt_number for display, because concurrent or resumed work doesn't always fit a neat counter.
At every external boundary you want a correlation ID. A tool call has a tool_call_id, an approval has an approval_id, and a change in a destination system has an action_id, which you use as the idempotency key where the destination supports one. Hang on to provider request IDs and destination object references too. They're what lets an operator line up your record with the logs in the other system.
Record both event time and recorded time. A slow queue can write an event well after the external action happened. Use a sequence number assigned by your task event store for ordering, and keep the destination's timestamps as what that system claims.
Task status should either be derived from events or updated under the same consistency rules. Something like running, waiting_for_approval, waiting_for_dependency, completed, completed_with_unresolved_action, cancelled or failed. Don't mark a task completed while an external action is only sitting in a queue.