27 February 2025 / Data systems / 8 chapters

Separate run state from item state

From Designing resumable large-data pipelines

A run says why and how a collection of items is being processed. An item says what happened to one unit. Putting both into a single status field makes partial recovery difficult because the run can fail while many items remain validly complete.

A run record might contain run_id, source version, configuration version, requested scope, creator, start time, stop reason and aggregate counts. Its lifecycle can stay small: prepared, running, pausing, paused, completed, completed_with_errors, cancelled and failed. The difference between paused and failed matters. A pause is an expected boundary with a known restart path. A failure says the controller could not continue under its current rules.

Item states need to describe recoverable facts rather than moods of the worker. A useful set is:

  • pending: known to the run but not claimed;
  • in_progress: claimed under a lease by a worker;
  • accepted: all required durable effects are confirmed;
  • rejected: input failed a defined validation rule;
  • retryable: an attempt failed and policy permits another;
  • quarantined: automated attempts have stopped pending review;
  • skipped: intentionally excluded with a recorded reason.

Keep attempt records separate when diagnosis matters. Each attempt can store its worker identity, lease token, start and finish times, stage, outcome and sanitised error. Overwriting one last_error field loses the sequence that explains whether a fault is consistent, transient or moving between stages.

State transitions should use compare-and-set conditions. A worker claiming a pending item updates it only if the state and lease are still what the worker observed. This prevents two workers from both believing they own the same item after a slow query or network delay.

All articles