Article chapter 03 of 08
Separate run state from item state
A run says why and how a collection of items is being processed. An item says what happened to one unit. Putting both into a single status field makes partial recovery difficult because the run can fail while many items remain validly complete.
A run record might contain run_id, source version, configuration version, requested scope, creator, start time, stop reason and aggregate counts. Its lifecycle can stay small: prepared, running, pausing, paused, completed, completed_with_errors, cancelled and failed. The difference between paused and failed matters. A pause is an expected boundary with a known restart path. A failure says the controller could not continue under its current rules.
Item states need to describe recoverable facts rather than moods of the worker. A useful set is:
pending: known to the run but not claimed;in_progress: claimed under a lease by a worker;accepted: all required durable effects are confirmed;rejected: input failed a defined validation rule;retryable: an attempt failed and policy permits another;quarantined: automated attempts have stopped pending review;skipped: intentionally excluded with a recorded reason.
Keep attempt records separate when diagnosis matters. Each attempt can store its worker identity, lease token, start and finish times, stage, outcome and sanitised error. Overwriting one last_error field loses the sequence that explains whether a fault is consistent, transient or moving between stages.
State transitions should use compare-and-set conditions. A worker claiming a pending item updates it only if the state and lease are still what the worker observed. This prevents two workers from both believing they own the same item after a slow query or network delay.