27 February 2025 / Data systems / 8 chapters

Keeping run state and item state apart

From Designing resumable large-data pipelines

A run describes why a set of items is being processed and under what settings. An item describes what happened to one unit. If you squash both into one status field, partial recovery gets hard, because a run can fail while plenty of its items are legitimately finished.

A run record might hold run_id, the source version, the configuration version, the requested scope, who created it, the start time, the stop reason and aggregate counts. Its lifecycle can stay small: prepared, running, pausing, paused, completed, completed_with_errors, cancelled and failed. Pay attention to the gap between paused and failed. A pause is an expected stopping point and you know how to restart from it. A failure means the controller couldn't keep going under its current rules.

Item states should describe facts you can recover from, rather than whatever mood the worker was in. Here's a set I'd start with:

  • pending: the run knows about it but nothing has claimed it
  • in_progress: a worker has claimed it under a lease
  • accepted: every required durable effect is confirmed
  • rejected: the input failed a defined validation rule
  • retryable: an attempt failed and the policy allows another one
  • quarantined: automatic attempts have stopped until someone reviews it
  • skipped: deliberately left out, with the reason recorded

If you care about diagnosing failures (and you will), keep attempt records in their own table. Each attempt can store the worker identity, lease token, start and finish times, the stage it reached, the outcome and a cleaned-up error. If you keep overwriting a single last_error field you lose the sequence, and the sequence is what tells you whether a fault is consistent, transient or moving between stages.

State changes should use compare-and-set conditions. When a worker claims a pending item, it only updates the row if the state and lease are still what it saw when it read them. That stops two workers from both thinking they own the same item after a slow query or a network delay.

All articles